Skip to content
#

agent-benchmarks

Here are 8 public repositories matching this topic...

Language: All
Filter by language
benchmark-radar

Track 10,000+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.

  • Updated Sep 7, 2026
  • Python
Awesome-AI4AI

AI4AI Survey: can AI reliably improve AI? 223 papers on long-horizon agents, benchmarks, harness design, and recursive self-improvement · updated weekly

  • Updated Sep 6, 2026
  • Python

Official companion repository for our survey "A Survey of the OpenClaw Ecosystem: From Platform Extensibility to Constraint Design" — a curated collection of papers, benchmarks, security reports, datasets, and tools for the OpenClaw AI agent ecosystem.

  • Updated May 31, 2026

用同一套提示词评测多模型 Agent 从零实现完整可玩游戏的能力:固定技术栈、固定评分标准、按模型归档对比。| A standardized prompt for benchmarking LLM agents on building a complete game from scratch — fixed tech stack, fixed rubric, per-model results.

  • Updated Aug 14, 2026
  • JavaScript

Add this topic to your repo

To associate your repository with the agent-benchmarks topic, visit your repo's landing page and select "manage topics."

Learn more