Human-in-the-loop adversarial workflows for high-stakes research audit: from ChatGPT-Gemini duels to 4-model MAD.
-
Updated
Jun 14, 2026
Human-in-the-loop adversarial workflows for high-stakes research audit: from ChatGPT-Gemini duels to 4-model MAD.
An adversarial AI expert workshop that stress-tests a research paper (rival-tradition referees argue; every comment quote-grounded and independently re-verified) and then rebuilds it: tracked-changes redline, clean version, your code re-run under a provenance wall, and a replication package. A Claude Code skill.
Three Claude Code skills for working with Codex CLI: codex-bridge (one-shot Codex calls), mad-build (Claude+Codex collaboration with cross-review), and mad-research (three-stream adversarial audit of papers, grants, reports with anonymized cross-critique and fresh-Codex synthesis).
A lightweight, portable agent skill for auditing research claims, evidence boundaries, and reproducibility readiness.
Faithful reproduction and reproducibility audit of Mall et al. (2022) - implemented twice (scikit-learn and pure NumPy), exposing 9-17 points of speaker leakage
A rigorous audit of Hybrid Quantum-Classical Networks (HQCNN) under noise and privacy constraints. (Outcome: Null Result / No Advantage Observed).
Falsification-first full-text novelty and public-artifact audit for version-aware agent memory
Audited reproduction and benchmark of Tencent/YOLO-Master Issue #50 with LoRA, strong baselines, multi-seed checks, protocol validation, and provenance tracking.
Add a description, image, and links to the research-audit topic page so that developers can more easily learn about it.
To associate your repository with the research-audit topic, visit your repo's landing page and select "manage topics."