Evidence-first tools and datasets for AI agents.
Making agent workflows easier to inspect, reproduce, and trust.
Research catalog · Evidence insights · Open-source projects
An auditable map of 2026 conference research that evaluates, analyzes, or outperforms Claude Code and Codex CLI as complete products—not papers that merely call a Claude or GPT model.
18,269 official records indexed · 13 product-level papers reviewed · Exact model and configuration evidence · Paper-linked insights
Explore the catalog · Read the insights · Use the dataset · Suggest a paper
| Project | What it does |
|---|---|
| Upstream Radar | Monitors agent-tooling dependencies for vulnerability, compatibility, and upstream changes. |
| Clean Your Data | Explains disk usage with bounded evidence and makes cleanup reversible. |
| WebMeld | Restyles webpages with natural language, verified previews, undo, and persistence. |
Evidence before claims · Local-first when practical · Reversible automation · Inspectability over magic
If the research catalog saves you time, give it a star ★—it helps more researchers find the evidence.

