bench: round 6 on Sonnet 5.5 — per-arm model, lean ADD arm; the Haiku advisor switch is refuted - #230
Merged
Merged
Conversation
…r; two lean arms switch models by beat Contract and red checks: arm tomls take model/advisor; run-all --model sets the model for arms without their own; argv carries --advisor only when set; records carry the resolved model; two arms (add-4-lean, add-4-advisor) share one lean skill variant that stubs less, greps personas and builds on a Haiku subagent. The arm count moves from 8 to 10. author: Tin Dang
…ess launch The autouse guard wrapped build_argv with a two-argument signature and raised whenever no agent was injected, so a test that replaces the launcher itself could not drive execute_wm. The guard moves to the launch layer, keeping its purpose: no test may start the real claude binary. Scope widens to benchmark/tests/conftest.py; R:NO_LIVE and C9 seal the safety property. author: Tin Dang
…ce's baseline commit C6 required the lean arms' first steps to equal all four add-4 steps, which would copy the variant skill after workspace_git.py commits the baseline and leave a modified SKILL.md in the working tree for the agent's Verify to trip on. C6 now requires every add-4 step in order, with the copy between the install and the baseline commit. author: Tin Dang
… arguments test_session_mode's spy took (prompt, agent_cmd) and test_wv2_family grepped core.py for the literal "model": PINNED_MODEL. Their intent is unchanged — every WM starts a fresh conversation, and every record stamps the model it ran on — so the spy passes new arguments through and the grep looks for the resolved model. Scope widens to both files. author: Tin Dang
… ADD arms switch models by beat An arm toml may now set `model` and `advisor`, and `run-all --model <id>` moves every arm that does not set its own. The live argv carries `--advisor` only when an arm sets one, and every record stamps the resolved model (and advisor), so a comparison can be checked for a mix-up. Without either, argv and records are unchanged. Two arms measure a per-beat model switch: add-4-lean installs the ADD 4.0 skill and overwrites it, before the workspace's baseline commit, with a variant that stubs only what the checks import, picks the lead persona by grep instead of reading the 65 KB index, and runs Build in one foreground Haiku subagent while the main session verifies. add-4-advisor runs the same variant on a Haiku 4.5 main session that consults Claude Sonnet 5.5 as its advisor. The no-live-agent test guard now refuses at process launch, so it holds however the argv is built. author: Tin Dang
Seal intact since the last refreeze; the checks pass fresh (10 passed) and the benchmark suite is green (530 passed, 12 skipped). Live round-6 records stamp each run's model and advisor, and the lean variant lands in the workspace's baseline commit with a clean tree. The harness works; what the models did with it is a measurement for the round-6 report: the Haiku main session never consulted its advisor, and the lean arm's Build stayed in the main session. author: Tin Dang
Round 6 on Claude Sonnet 5.5 refuted the model switch. The Haiku main session never consulted its Sonnet advisor in 6 runs, failed all 3 wm1 apps on Flask/FastAPI, and cost 1.6-1.8x plain add-4 at twice the wall time; the lean arm never handed Build to a Haiku subagent. The contract retires the advisor arm (10 -> 9 arms), cuts the lean variant back to its two work-cutting edits, and seals that any arm pinning a model pins a Sonnet. The harness's model and advisor keys stay, since `--model claude-sonnet-5-5` needs them. Four checks fail red. author: Tin Dang
…ps Build in the Sonnet session The add-4-advisor arm and its name leave the roster (10 -> 9 arms). The lean variant drops its hand-off of Build to a Haiku subagent and returns to the shipped Build line, so it differs from the shipped skill only where it cuts work: stub only what the checks import, and pick the lead persona by grep instead of reading the 65 KB index. The model and advisor arm keys stay for `run-all --model`. author: Tin Dang
Seal intact; 11 checks pass fresh, 531 benchmark tests pass. The lean variant differs from the shipped skill in exactly two places and names no Haiku model; nine arms remain. author: Tin Dang
…s by beat is refuted On claude-sonnet-5-5 the same ADD 4.0 skill ran Direction in 1.1-1.7 min against 4.1 on Sonnet 5, and cost 1.7-2.1x vanilla against 2.8-2.9x. A Haiku main session with a Sonnet advisor never consulted it in 6 runs, cost 1.6-1.8x add-4, and failed every wm1 app on a non-stdlib framework; the lean skill never handed Build to a Haiku subagent. The lean edits that cut work (fewer stubs, persona by grep) trimmed wall time 14-38% with no quality loss. On this model both workloads saturate on edges and mutation; ADD keeps its lead on the ambiguous spec. n = 3, not significant. author: Tin Dang
| from benchmark.check_isolation import find_leaks | ||
| from benchmark.check_isolation import main as check_isolation_main | ||
| from benchmark.runner.agent import PINNED_MODEL, build_argv | ||
| from benchmark.runner.agent import PINNED_MODEL, build_argv, resolve_model |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Round 6 of the ADD 4.0 pilot, on
claude-sonnet-5-5at medium effort: does a newer model, a leaner skill, or a cheaper main model with a stronger advisor cut Direction's time and cost?model/advisor, andrun-all --model <id>moves every arm that doesn't set its own. Records stamp the resolved model (and advisor). With neither set, argv and records are unchanged (claude-sonnet-5).add-4-leanruns a two-edit variant of the shipped skill (stub only what the checks import; pick the persona by grep). The Haiku-mainadd-4-advisorarm was measured, refuted and retired, so the roster is 9 arms.benchmark/PILOT-4v3-2026-09-30-r6.md.Findings (n = 3 per cell, not significant)
ADD tasks
bench-model-switch: freeze, 3 refreezes, verify PASSbench-sonnet-only: freeze, verify PASSTest plan
python3 -m pytest -q benchmark/tests/test_arm_models.py benchmark/tests/test_arms.py: 11 passedpython3 -m pytest -q benchmark/tests: 531 passed, 12 skipped