Skip to content

bench: round 6 on Sonnet 5.5 — per-arm model, lean ADD arm; the Haiku advisor switch is refuted - #230

Merged
TinDang97 merged 10 commits into
mainfrom
feat/bench-model-switch
Sep 30, 2026
Merged

TinDang97 merged 10 commits into
mainfrom
feat/bench-model-switch

Conversation

@TinDang97

Copy link
Copy Markdown
Collaborator

Summary

Round 6 of the ADD 4.0 pilot, on claude-sonnet-5-5 at medium effort: does a newer model, a leaner skill, or a cheaper main model with a stronger advisor cut Direction's time and cost?

  • Harness: an arm toml may set model / advisor, and run-all --model <id> moves every arm that doesn't set its own. Records stamp the resolved model (and advisor). With neither set, argv and records are unchanged (claude-sonnet-5).
  • Guard: the no-live-agent test guard now refuses at process launch, so it holds however the argv is built.
  • Arms: add-4-lean runs a two-edit variant of the shipped skill (stub only what the checks import; pick the persona by grep). The Haiku-main add-4-advisor arm was measured, refuted and retired, so the roster is 9 arms.
  • Report: benchmark/PILOT-4v3-2026-09-30-r6.md.

Findings (n = 3 per cell, not significant)

  • Sonnet 5.5 alone cut Direction from 4.1 min (round-5 mean) to 1.1 min on wm1 and 1.7 min on amb1. ADD's cost went from $1.67 to $0.52 on wm1 (1.7× vanilla, down from 2.8×).
  • The Haiku main session never consulted its Sonnet advisor in 6 runs. It cost 1.6–1.8× add-4, and all 3 wm1 apps were built on Flask/FastAPI, which fail the stdlib-only entry contract.
  • The lean edits cut wall time 14–38% with no quality loss.
  • On Sonnet 5.5 both workloads saturate on edges and mutation. ADD keeps its lead on the ambiguous spec (5.7 vs 4.3 of 7).

ADD tasks

  • bench-model-switch: freeze, 3 refreezes, verify PASS
  • bench-sonnet-only: freeze, verify PASS

Test plan

  • python3 -m pytest -q benchmark/tests/test_arm_models.py benchmark/tests/test_arms.py: 11 passed
  • python3 -m pytest -q benchmark/tests: 531 passed, 12 skipped

…r; two lean arms switch models by beat

Contract and red checks: arm tomls take model/advisor; run-all --model sets the model for arms
without their own; argv carries --advisor only when set; records carry the resolved model; two
arms (add-4-lean, add-4-advisor) share one lean skill variant that stubs less, greps personas and
builds on a Haiku subagent. The arm count moves from 8 to 10.

author: Tin Dang
…ess launch

The autouse guard wrapped build_argv with a two-argument signature and raised whenever no agent
was injected, so a test that replaces the launcher itself could not drive execute_wm. The guard
moves to the launch layer, keeping its purpose: no test may start the real claude binary. Scope
widens to benchmark/tests/conftest.py; R:NO_LIVE and C9 seal the safety property.

author: Tin Dang
…ce's baseline commit

C6 required the lean arms' first steps to equal all four add-4 steps, which would copy the variant
skill after workspace_git.py commits the baseline and leave a modified SKILL.md in the working tree
for the agent's Verify to trip on. C6 now requires every add-4 step in order, with the copy between
the install and the baseline commit.

author: Tin Dang
… arguments

test_session_mode's spy took (prompt, agent_cmd) and test_wv2_family grepped core.py for the
literal "model": PINNED_MODEL. Their intent is unchanged — every WM starts a fresh conversation,
and every record stamps the model it ran on — so the spy passes new arguments through and the
grep looks for the resolved model. Scope widens to both files.

author: Tin Dang
… ADD arms switch models by beat

An arm toml may now set `model` and `advisor`, and `run-all --model <id>` moves every arm that
does not set its own. The live argv carries `--advisor` only when an arm sets one, and every
record stamps the resolved model (and advisor), so a comparison can be checked for a mix-up.
Without either, argv and records are unchanged.

Two arms measure a per-beat model switch: add-4-lean installs the ADD 4.0 skill and overwrites it,
before the workspace's baseline commit, with a variant that stubs only what the checks import,
picks the lead persona by grep instead of reading the 65 KB index, and runs Build in one
foreground Haiku subagent while the main session verifies. add-4-advisor runs the same variant on
a Haiku 4.5 main session that consults Claude Sonnet 5.5 as its advisor.

The no-live-agent test guard now refuses at process launch, so it holds however the argv is built.

author: Tin Dang
Seal intact since the last refreeze; the checks pass fresh (10 passed) and the benchmark suite
is green (530 passed, 12 skipped). Live round-6 records stamp each run's model and advisor, and
the lean variant lands in the workspace's baseline commit with a clean tree.

The harness works; what the models did with it is a measurement for the round-6 report: the
Haiku main session never consulted its advisor, and the lean arm's Build stayed in the main session.

author: Tin Dang
Round 6 on Claude Sonnet 5.5 refuted the model switch. The Haiku main session never consulted
its Sonnet advisor in 6 runs, failed all 3 wm1 apps on Flask/FastAPI, and cost 1.6-1.8x plain
add-4 at twice the wall time; the lean arm never handed Build to a Haiku subagent.

The contract retires the advisor arm (10 -> 9 arms), cuts the lean variant back to its two
work-cutting edits, and seals that any arm pinning a model pins a Sonnet. The harness's model
and advisor keys stay, since `--model claude-sonnet-5-5` needs them. Four checks fail red.

author: Tin Dang
…ps Build in the Sonnet session

The add-4-advisor arm and its name leave the roster (10 -> 9 arms). The lean variant drops its
hand-off of Build to a Haiku subagent and returns to the shipped Build line, so it differs from
the shipped skill only where it cuts work: stub only what the checks import, and pick the lead
persona by grep instead of reading the 65 KB index. The model and advisor arm keys stay for
`run-all --model`.

author: Tin Dang
Seal intact; 11 checks pass fresh, 531 benchmark tests pass. The lean variant differs from the
shipped skill in exactly two places and names no Haiku model; nine arms remain.

author: Tin Dang
…s by beat is refuted

On claude-sonnet-5-5 the same ADD 4.0 skill ran Direction in 1.1-1.7 min against 4.1 on Sonnet 5,
and cost 1.7-2.1x vanilla against 2.8-2.9x. A Haiku main session with a Sonnet advisor never
consulted it in 6 runs, cost 1.6-1.8x add-4, and failed every wm1 app on a non-stdlib framework;
the lean skill never handed Build to a Haiku subagent. The lean edits that cut work (fewer stubs,
persona by grep) trimmed wall time 14-38% with no quality loss. On this model both workloads
saturate on edges and mutation; ADD keeps its lead on the ambiguous spec. n = 3, not significant.

author: Tin Dang
Comment thread benchmark/runner/core.py
from benchmark.check_isolation import find_leaks
from benchmark.check_isolation import main as check_isolation_main
from benchmark.runner.agent import PINNED_MODEL, build_argv
from benchmark.runner.agent import PINNED_MODEL, build_argv, resolve_model
@TinDang97
TinDang97 merged commit 8817f53 into main Sep 30, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant