Track 10,000+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
-
Updated
Sep 7, 2026 - Python
Track 10,000+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
AI4AI Survey: can AI reliably improve AI? 223 papers on long-horizon agents, benchmarks, harness design, and recursive self-improvement · updated weekly
Official companion repository for our survey "A Survey of the OpenClaw Ecosystem: From Platform Extensibility to Constraint Design" — a curated collection of papers, benchmarks, security reports, datasets, and tools for the OpenClaw AI agent ecosystem.
An evidence-hostile, container-isolated benchmark and behavior analysis platform for long-horizon AI software agents.
A free Claude Code skill to test coding agents on the same tasks; compare pass rate, cost, and time—built on affaan-m/ECC.
Generated by Codex. Read with caution.
Deterministic release decisions from agent benchmark evidence.
用同一套提示词评测多模型 Agent 从零实现完整可玩游戏的能力:固定技术栈、固定评分标准、按模型归档对比。| A standardized prompt for benchmarking LLM agents on building a complete game from scratch — fixed tech stack, fixed rubric, per-model results.
To associate your repository with the agent-benchmarks topic, visit your repo's landing page and select "manage topics."