Skip to content

Repository files navigation

Coding-agent experiments flowing through an evidence graph into research papers

Awesome Claude Code & Codex Papers

Research that measures, analyzes, or beats real production coding agents.

English · 简体中文

Open website Compare methods Read insights

Validate catalog GitHub stars MIT License PRs welcome Data YAML

papers: 23 official records indexed: 20,673 direct comparisons: 10 official artifacts: 18 domains: 8 conference series tracked: 13 reviewed: 2026-08-25

Coverage at a glance

Research domains
Software Engineering · 14 · Security · 4 · Systems & Performance · 2 · Machine Learning · 5 · Scientific Computing · 5 · Formal Methods · 4 · Web & UI · 3 · Documents · 1

Conferences: catalog / official records
AAAI · 0 / 4,920 · ASE · 1 / 263 · FSE · 0 / 211 · ICLR · 12 / 5,351 · ICML · 5 / 6,628 · ICSE · 0 / 321 · ISSTA · 3 / 210 · NeurIPS · list pending · IJCAI · 1 / 989 · KDD · 1 / 1,415 · PLDI · 0 / 106 · POPL · 0 / 92 · OOPSLA · 0 / 167

This repository is the open data and maintenance layer behind the web-first research catalog. The main catalog is restricted to 2026 papers with an official conference, proceedings, or OpenReview conference record that evaluate, analyze, or outperform Claude Code and Codex CLI as complete products—not papers that merely use a Claude or GPT-family model. Every official-list record remains auditable as included, excluded, pending, or duplicate, so a compact catalog reflects strict scope rather than silent filtering. Official records prove venue identity; identity-verified open copies may supply content evidence but never replace the primary paper URL.

Start with the website

The interactive catalog is the fastest way to:

  • search systems, tasks, methods, authors, and reported results;
  • choose a research domain, conference, year, product, evidence strength, exact model, or method;
  • inspect baseline models, versions, budgets, evidence locations, and caveats;
  • switch between English and Chinese without reading giant Markdown tables.

The separate evidence insights page synthesizes where the products struggle and maps every inference back to the supporting papers, reported results, source locations, and comparison caveats.

Use the methods matrix to compare interventions without separating results from their controls, or reuse the official-conference census skill to build another auditable venue-wide scan.

Evidence standard

This is not a leaderboard. Product versions, backbone models, budgets, tools, and task domains often differ. Every entry separates the paper's reported result from our comparison controls and caveats. Missing details stay explicitly unknown.

For raw or generated research material, use the paper dossiers, domain view, conference view, machine-readable JSON, or the full 2026 conference census.

Contributing

Suggest a paper, correct an evidence record, join a research discussion, or read CONTRIBUTING.md. The catalog is generated from data/papers.yaml and validated in CI.


Did this save you a paper-reading session?
Star the repository ★ — it helps more researchers find the evidence.

Releases

Packages

Used by

Contributors

Languages