Two halves. Below, the work I built and the work I landed in other people's projects. Every pull request in the second half is a commit that shipped upstream: a defect I reproduced first, the smallest fix I could defend, and the tests that pin the behaviour. Both halves are rendered from the GitHub API on a schedule, so neither can drift from what is actually public.
Focus · AI agent runtimes and their memory subsystems · multi-agent orchestration · MCP and tool plumbing · multimodal agents · Python
   
These 16 labs were each built from scratch and run on a laptop CPU. Every one is deterministic, offline and CI-tested, and every figure its README reports is re-derived by a test - from a committed artifact, or by re-running the CLI and byte-comparing the transcript - so no headline number in any of them is hand-typed; the measurement labs run several seeds and print the spread instead of a best run. The descriptions below are pulled live from each repository, so this table cannot rot either.
| Repository | What it demonstrates |
|---|---|
| align-lab CPU-only post-training study: SFT vs DPO vs ORPO vs SimPO on a verifiable digit-addition task. Multi-seed; shows DPO/SimPO win-rate climbing to ~0.99 while real generation accuracy collapses (likelihood displacement). |
Four preference-optimisation objectives on one task and one policy: the pair-wise win-rate is a flattering metric that can rise while generation gets worse. |
| grpo-repro A tiny, fully-verifiable GRPO / REINFORCE / DPO comparison lab on a Reverse-Polish-Notation puzzle env. Same policy, same reward, same budget; CPU-measured with per-run wall-clock reported honestly. |
GRPO, REINFORCE and DPO as three arms on identical weights and an identical verifiable reward. With a partial credit given to any legal answer, all three collapse to a degenerate output; the reward shape, not the algorithm, was the bug. |
| reward-hacking-lab CPU-only study of reward-model over-optimisation (Goodharting) in RLHF-style rejection-sampling self-improvement: optimise a verifiable reward and accuracy climbs; optimise a learned reward and the proxy climbs while true accuracy collapses. |
Same loop, two rewards: optimise an exact verifier and true accuracy climbs, optimise a learned proxy and the proxy climbs while accuracy falls. Goodhart on demand, measured. |
| starlab CPU-reproducible STaR self-improvement: a ~100k-param transformer bootstraps column-addition reasoning from its own verifier-checked chains, with a matched-compute control. |
STaR bootstrap: the model samples its own chains, an exact arithmetic verifier filters them, survivors become next round's training set - reported against a matched-compute control so the gain cannot be mistaken for simply training longer. |
| data-select-lab A CPU-only LoRA/PEFT + LESS data-selection study: influence-scored LoRA gradient features pick the few fine-tuning examples that teach a frozen model a held-out capability, with forgetting reported honestly. |
LESS-style LoRA gradient features pick the few dozen examples that teach a frozen model a held-out skill; the same table reports how much of the skill it already had that costs. |
| Repository | What it demonstrates |
|---|---|
| skill-lab CPU-only, dependency-free study of a Voyager-style self-evolving skill library: a budgeted BFS planner distils reusable first-order skills from its own verified plans and gets measurably cheaper the more it is used. |
Voyager-style skill library distilled from the agent's own verified plans: the budget spent on the next task falls as the library grows, which is the whole claim. |
| autprompt-lab CPU-only study of budgeted automatic prompt optimisation (OPRO / Promptbreeder-style): random vs greedy vs population-evolution search against a frozen tiny transformer on a verifier-scored task, at an equal query budget. |
OPRO / Promptbreeder-style prompt search under an equal query budget - random, greedy and population evolution scored against one frozen policy, so the comparison is about search, not compute. |
| evolve-lab A CPU-only, dependency-free self-optimizing agent: evolution rediscovers and blends single-machine dispatching rules purely from an exact tardiness verifier. |
A (mu, lambda) evolution strategy over dispatch rules rediscovers WSPT from an exact tardiness verifier. Dependency-free, and honest that the margin it keeps over WSPT sits inside the seed noise. |
| Repository | What it demonstrates |
|---|---|
| lagent Local-first ReAct tool-use agent with a measured scaffold ablation: identical behaviour-cloned weights, four harnesses, ground-truth verifier. CPU-only tiny transformer. |
ReAct scaffold ablation on identical behaviour-cloned weights: strip the observation scratchpad and the solve rate collapses. The harness loop is load-bearing machinery, not decoration around the model. |
| agent-harness-eval Agent Harness Eval: instrument, score, and statistically compare agent rollouts. Paired bootstrap + McNemar, behavior metrics, zero-dependency core. |
Paired bootstrap and McNemar for comparing agent rollouts, because a three-seed difference is not a result and most harness comparisons never check. |
| hier-memo-agents Hierarchical planner-executor agents with shared blackboard memory, topic-level reuse, and verifiable citation provenance. Zero deps, deterministic, CI-tested. |
Planner and executor agents over a shared blackboard with citation-linked reuse: memory matched by topic rather than by question, so reuse is transfer and not a cache hit. |
| research-relay-lab CPU-only, multi-seed measurement of where a deep-research pipeline loses the evidence it found: one frozen 142k-parameter policy run through six researcher-to-reporter harnesses on a task with an exact verifier, with every README number rendered from a committed artifact. |
A deer-flow-shaped Planner / Research Team / Reporter pipeline measured at the step nobody evaluates: the reporter handed raw search hits loses no facts at all and still lands near chance, and one notebook slot short of the chain every report cites more hops than its own notebook holds. |
| Repository | What it demonstrates |
|---|---|
| llama-anatomy From-scratch ~100k-parameter LLaMA decoder (RMSNorm, RoPE, SwiGLU, GQA), plus an equal-parameter ablation of each choice, a train-short/test-long probe of RoPE against learned absolute positions, and a parameter-free sweep of RoPE's rotation base. CPU-only, multi-seed; every README number renders from a committed artifact and CI byte-compares it. |
A LLaMA decoder written from the papers at 10^5 parameters, with every component swapped for an equal-parameter control so a difference is never just size. The rotation-base sweep is the one axis that adds no parameters at all, and it separates a floor its seeds agree on from a ceiling they do not. |
| consistency-lab Adaptive early-stopping self-consistency: a sequential stopping rule that halts CoT sampling once the vote is statistically decided, measured on the accuracy-vs-compute frontier vs fixed-N self-consistency. CPU-only, verifiable. |
A sequential stopping rule for self-consistency: sample chains until the vote is statistically decided, then spend the chains you saved on the questions that are actually contested. |
| vlm-distill-bench Reproducible CPU-only distillation & quantization benchmark for compact vision-language models on procedural mini-CLEVR. Real numbers, committed seeds, no downloads. |
Distillation plus quantisation of a compact vision-language model on procedural mini-CLEVR, including the negative result that temperature-scaled KD can lose to plain cross-entropy at this scale. |
| system-one-lab CPU-only, multi-seed measurement of what an agent decision layer is sensitive to: a typed Choice/Score/Noul judge vs a prose judge vs a prompt-flattened control, on synthetic findings with an exact verifier. |
Three judges over the same findings, separated only by what they can read: the control that is handed the evidence in its prompt still changes its answer when only the wording moves, and one global temperature turn makes the pooled error smaller by making the honest slice worse. |
Repositories run by a company or product organisation, each with 10,000+ stars, that have merged my pull requests upstream.
| Project | Stars | Merged | Pull requests |
|---|---|---|---|
| bytedance/deer-flow | 83.3K | 14 | #5555 · #5593 · #5609 · #5586 · #5588 · #5801 · #5821 · #5607 · #5852 · #5864 · #5883 · #5850 · #5960 · #6067 |
| agentscope-ai/agentscope | 32.6K | 3 | #2754 · #2808 · #2805 |
| deepset-ai/haystack | 26.6K | 3 | #12810 · #12905 · #12907 |
| modelscope/ms-swift | 15.8K | 5 | #10245 · #10248 · #10247 · #10250 · #10246 |
| livekit/agents | 14.4K | 1 | #7468 |
| The-PR-Agent/pr-agent | 13.2K | 1 | #3722 |
Merged into community and individual-maintainer projects (28 pull request(s) in 5 project(s))
Also 10,000+ stars, but owned by a solo maintainer or an academic / community group rather than a company.
| Project | Stars | Merged | Pull requests |
|---|---|---|---|
| zhayujie/CowAgent | 47.2K | 19 | #3247 · #3248 · #3284 · #3282 · #3289 · #3291 · #3293 · #3299 · #3296 · #3303 · #3305 · #3312 · #3314 · #3316 · #3321 · #3319 · #3325 · #3332 · #3344 |
| AstrBotDevs/AstrBot | 41.3K | 3 | #10229 · #10231 · #10248 |
| andrewyng/openworker | 18.4K | 1 | #690 |
| StarTrail-org/LEANN | 13.0K | 1 | #427 |
| mrexodia/ida-pro-mcp | 12.4K | 4 | #531 · #529 · #533 · #539 |
Merged into projects under 10,000 stars (4 pull request(s) in 3 project(s))
Real merges, below the 10,000 star bar of the tables above.
| Project | Stars | Merged | Pull requests |
|---|---|---|---|
| TencentCloud/Octop | 6.2K | 1 | #1150 |
| QwenLM/Qwen-MM-Plugins | 3.1K | 2 | #69 · #71 |
| iot-hackathon-2026-teamESGenius/iot_hackathon_2026_mannings | 0 | 1 | #1 |
Merged per month
2026-02 # 1
2026-09 ######################################## 53
2026-10 #### 5
All merged pull requests
| Merged | Project | Pull request |
|---|---|---|
| 2026-10-01 | QwenLM/Qwen-MM-Plugins | #71 fix(api): bound transcribe_audio's ffmpeg calls with QWEN_MM_FFMPEG_TIMEOUT |
| 2026-10-01 | QwenLM/Qwen-MM-Plugins | #69 fix(blender,freecad): keep a blank port in the config from killing the server |
| 2026-10-01 | zhayujie/CowAgent | #3344 fix(env_config): refuse entries that would split into extra .env lines |
| 2026-10-01 | zhayujie/CowAgent | #3332 fix(knowledge): keep graph node paths slash-separated on Windows |
| 2026-10-01 | bytedance/deer-flow | #6067 fix(channels): split Telegram messages by UTF-16 code units, not code points |
| 2026-09-30 | bytedance/deer-flow | #5960 fix(agents): anchor agent-name validation against a trailing newline |
| 2026-09-30 | bytedance/deer-flow | #5850 fix(config): guard the recovered stream cleanup delay like its heartbeat sibling |
| 2026-09-30 | deepset-ai/haystack | #12907 fix: render every content part of ChatPromptBuilder message templates |
| 2026-09-30 | andrewyng/openworker | #690 fix(mentions): a corrupt mention_threads.json must not brick server startup |
| 2026-09-30 | livekit/agents | #7468 fix(voice): treat a non-list record_keyterms argument as no change |
| 2026-09-29 | StarTrail-org/LEANN | #427 fix(core): honour an explicit passage ID of 0 and key the offset map by string |
| 2026-09-29 | zhayujie/CowAgent | #3325 fix(cli): remove the temp dir when a repo archive cannot be extracted |
| 2026-09-29 | zhayujie/CowAgent | #3319 fix(cli): overlay the live roster when restoring a backup |
| 2026-09-29 | zhayujie/CowAgent | #3321 fix(wechatmp): upload a local video reply instead of crashing on its path |
| 2026-09-29 | The-PR-Agent/pr-agent | #3722 fix(azure): do not resolve a comment thread after a swallowed tool failure |
| 2026-09-29 | zhayujie/CowAgent | #3316 fix(channel): cancel queued work under the agent-scoped session key |
| 2026-09-29 | zhayujie/CowAgent | #3314 fix(plugins): normalize a damaged plugins.json entry instead of losing every plugin |
| 2026-09-29 | zhayujie/CowAgent | #3312 fix(wechatcom): read a locally generated image instead of fetching it |
| 2026-09-29 | zhayujie/CowAgent | #3305 fix(godcmd): answer when #model is given more than one argument |
| 2026-09-29 | zhayujie/CowAgent | #3303 fix(godcmd): answer when the plugin priority is not a number |
| 2026-09-29 | zhayujie/CowAgent | #3296 fix(role): answer when the customize command has no description |
| 2026-09-29 | zhayujie/CowAgent | #3299 fix(feishu,dingtalk): upload local video replies instead of dropping them |
| 2026-09-29 | agentscope-ai/agentscope | #2805 fix(message): keep a copy of the usage appended to a message |
| 2026-09-29 | zhayujie/CowAgent | #3293 fix(dingtalk): route text-only richText to the agent instead of dropping it |
| 2026-09-29 | zhayujie/CowAgent | #3291 fix(linkai): reply when the midjourney image index is not a number |
| 2026-09-29 | zhayujie/CowAgent | #3289 fix(discord): upload local video replies instead of printing the file:// path |
| 2026-09-29 | zhayujie/CowAgent | #3282 fix(dingtalk): deliver ERROR and INFO replies instead of dropping them |
| 2026-09-29 | zhayujie/CowAgent | #3284 fix(dingtalk): skip group messages outside the whitelist instead of crashing |
| 2026-09-28 | modelscope/ms-swift | #10246 fix(utils): treat a blank or non-numeric LOCAL_RANK as unset |
| 2026-09-28 | modelscope/ms-swift | #10250 fix(infer): raise the intended ValueError when tool_choice names an unknown tool |
| 2026-09-28 | bytedance/deer-flow | #5883 fix(community): normalize InfoQuest timeout config values like the sibling providers do |
| 2026-09-28 | zhayujie/CowAgent | #3248 fix(wechat_kf): keep a server-supplied file name inside the tmp dir |
| 2026-09-28 | zhayujie/CowAgent | #3247 fix(telegram): keep a sender-chosen document name inside the tmp dir |
| 2026-09-27 | AstrBotDevs/AstrBot | #10248 fix(platform): keep the Satori heartbeat and reconnect delays from being disabled by a cleared field |
| 2026-09-27 | bytedance/deer-flow | #5864 fix(models): skip credential files whose access token is not a string |
| 2026-09-27 | AstrBotDevs/AstrBot | #10231 fix(respond): keep an unusable segmented-reply log base from dropping the reply |
| 2026-09-27 | AstrBotDevs/AstrBot | #10229 fix(weixin_oc): keep a cleared timeout field from disabling the HTTP timeout |
| 2026-09-26 | modelscope/ms-swift | #10247 fix(metrics): score only sequences that carry a supervised label in seq_acc |
| 2026-09-26 | modelscope/ms-swift | #10248 fix(cli): stringify list values from a YAML/JSON config before building argv |
| 2026-09-26 | TencentCloud/Octop | #1150 fix(config): bound OCTOP_PORT and --port to a bindable range |
| 2026-09-26 | modelscope/ms-swift | #10245 fix(utils): fall back to INFO for a blank or unknown LOG_LEVEL |
| 2026-09-25 | bytedance/deer-flow | #5852 fix(community): reject bool and fractional web-search max_results like image search does |
| 2026-09-24 | bytedance/deer-flow | #5607 fix(memory): reject a Honcho base_url that can never resolve |
| 2026-09-24 | deepset-ai/haystack | #12905 fix(core): compare Ellipsis callable parameters by return type |
| 2026-09-24 | bytedance/deer-flow | #5821 fix(community): fall back to the default SearXNG max_results on an unparseable value |
| 2026-09-24 | bytedance/deer-flow | #5801 fix(gateway): treat a blank GATEWAY_HOST/GATEWAY_PORT as unset |
| 2026-09-24 | bytedance/deer-flow | #5588 fix(skills): load SKILL.md saved as UTF-8 with a BOM |
| 2026-09-24 | bytedance/deer-flow | #5586 fix(memory): keep a zero confidence from scoring as the default |
| 2026-09-24 | agentscope-ai/agentscope | #2808 fix(formatter): forward the images of collapsed messages in xAI multi-agent history |
| 2026-09-23 | deepset-ai/haystack | #12810 fix(auth): deserialize listed secrets when recursive is enabled |
| 2026-09-22 | mrexodia/ida-pro-mcp | #539 fix(rpc): treat a blank IDA_MCP_URL as unset for download URLs |
| 2026-09-22 | bytedance/deer-flow | #5609 fix(subagents): report an explicit zero batch limit instead of defaulting it |
| 2026-09-22 | bytedance/deer-flow | #5593 fix(skills): render an explicit empty allowed-tools as no tools, not as all |
| 2026-09-22 | agentscope-ai/agentscope | #2754 fix(agent): keep tool-result metadata and timestamps when truncating |
| 2026-09-21 | mrexodia/ida-pro-mcp | #533 fix(idalib): fall back to defaults for blank IDA_MCP_* values |
| 2026-09-21 | mrexodia/ida-pro-mcp | #529 fix(installer): keep IPv6 brackets in generated transport URLs |
| 2026-09-21 | mrexodia/ida-pro-mcp | #531 fix(installer): let --uninstall run when IDA Free is installed |
| 2026-09-20 | bytedance/deer-flow | #5555 fix(memory): report malformed backend_config values by key name |
| 2026-02-12 | iot-hackathon-2026-teamESGenius/iot_hackathon_2026_mannings | #1 feat(routing): enhance robust optimizer with learning strategy |
In review right now (48 open pull requests)
| Project | Open | Pull requests |
|---|---|---|
| agno-agi/agno | 7 | #10281 · #10287 · #10297 · #10301 · #10320 · #10323 · #10325 |
| browser-use/browser-use | 4 | #5838 · #5845 · #5847 · #5849 |
| crewAIInc/crewAI | 4 | #7610 · #7612 · #7615 · #7617 |
| langflow-ai/langflow | 4 | #15223 · #15225 · #15227 · #15229 |
| mrexodia/ida-pro-mcp | 4 | #537 · #541 · #542 · #543 |
| TencentCloud/Octop | 4 | #910 · #912 · #1146 · #1147 |
| modelscope/ms-swift | 3 | #10251 · #10264 · #10266 |
| strands-agents/harness-sdk | 3 | #4553 · #4554 · #4591 |
| agentscope-ai/agentscope | 2 | #2741 · #2799 |
| HKUDS/LightRAG | 2 | #4097 · #4098 |
| HKUDS/nanobot | 2 | #5913 · #5914 |
| run-llama/llama_index | 2 | #23244 · #23248 |
| AstrBotDevs/AstrBot | 1 | #10286 |
| confident-ai/deepeval | 1 | #3377 |
| deepset-ai/haystack | 1 | #12813 |
| livekit/agents | 1 | #7470 |
| microsoft/agent-framework | 1 | #8787 |
| PrefectHQ/prefect | 1 | #23190 |
| StarTrail-org/LEANN | 1 | #428 |
Rendered by scripts/build_readme.py from the GitHub API — the lab table reads each repository's live description, the PR tables read the search API — and refreshed by .github/workflows/refresh.yml · record last changed 2026-10-02 00:45 UTC.

