Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
60 changes: 60 additions & 0 deletions .add/tasks/bench-model-switch.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
---
type: Task
title: the benchmark can run each arm on its own model and advisor, with two lean ADD arms that switch models by beat
status: done
kind: feature
risks: [measurement-validity, compatibility]
scope: [benchmark/arms/, benchmark/runner/agent.py, benchmark/runner/core.py, benchmark/pilot.py, benchmark/tests/test_arm_models.py, benchmark/tests/test_arms.py, benchmark/tests/conftest.py, benchmark/tests/test_session_mode.py, benchmark/tests/test_wv2_family.py]
gives: [S1 `run-all --model <id>` and the arm-toml keys `model` / `advisor`]
---
## CARD
goal: run the round-6 benchmark on claude-sonnet-5-5 and compare two model-switching ADD arms — Build on a Haiku subagent, and a Haiku main session with a Sonnet 5.5 advisor — against ADD and vanilla
why: Direction is 56% of ADD's wall time (4.1 of 7.3 min) and about 44% of its tokens; the harness pins every arm to claude-sonnet-5, so neither a newer model nor a per-beat model switch can be measured

## RULES
- M1 an arm toml may set `model` and `advisor`; both default to empty (from: request — "smart model switcher by advisor")
- M2 `run-all --model <id>` (and `run_reps` / `run_pilot` `model=`) sets the model for every arm that does not set its own; the resolved model is `arm.model` or the `--model` value or `PINNED_MODEL` (from: request — "run benchmark use sonnet-5-5")
- M3 the live argv carries `--model <resolved>` and `--effort medium`, plus `--advisor <advisor>` only when the arm sets one (from: research — `--advisor <model>` works headless, probed 2026-09-30)
- M4 every record's artifacts carry the resolved model, and `advisor` when set, so a comparison can be checked for a model mix-up (from: benchmark/runner/agent.py — the model-pin rationale)
- M5 two arms load: `add-4-lean` runs every add-4 setup step and overwrites the installed SKILL.md with `benchmark/arms/variants/add-4-lean/SKILL.md` after the install and before the workspace's baseline commit; `add-4-advisor` does the same with `model = "claude-haiku-4-5-20251001"` and `advisor = "claude-sonnet-5-5"` (from: request) (derived: the two arms share one variant, so the model switch is the only difference between them)
- M6 the lean variant differs from the shipped skill in exactly three places: stubs only what the checks import; the lead persona is picked by grepping `.add/personas/`, the 65 KB index only grepped when none fits; Build runs in one foreground subagent on `model: haiku`, and the main session verifies (from: PILOT r5 transcripts — 10.7 Direction writes per run, about half stubs; 5 of 6 runs opened the 65 KB index; Build is 53–54% of tokens)
- R:DEFAULT with no `--model` and no arm model, argv and records are exactly as before — `claude-sonnet-5` (from: benchmark/tests — the model-pin tests)
- R:NO_LIVE no test launches the real `claude` binary; the guard refuses at process launch, so a test that replaces the launcher itself may drive `execute_wm` without an injected agent (from: benchmark/tests/conftest.py — the 2026-09-28 live-spend incident)
- R:ARMCOUNT `ARM_NAMES` grows from 8 to 10; the fairness fields stay identical across all arms (from: benchmark/tests/test_arms.py::test_all_arms_validate_with_fairness_parity — its count changes with this contract)

## ASSUMPTIONS
- A1 [which] "sonnet-5-5" means `claude-sonnet-5-5` → found: the model id answers; the bare alias `sonnet-5-5` is refused (evidence: `claude -p … --model claude-sonnet-5-5` → OK; `--model sonnet-5-5` → unrecognized_model)
- A2 [which] Haiku 4.5 takes `--effort medium` and `--advisor` together → found: it does (evidence: `claude -p "reply with just OK" --model claude-haiku-4-5-20251001 --effort medium --advisor claude-sonnet-5-5` → OK)
- A3 [experience] an advisor call is expensive — it received 57.7k uncached input tokens in one consult ($0.13 of a $0.18 run) → the advisor arm may cost more than it saves if Haiku consults often → measured, not assumed
- A4 [which] the variant is measured before it is shipped → it lives under benchmark/arms/variants/, not in add-method/skill/
- A5 [who] a Haiku Build subagent may write weaker code → the sealed checks and the main-session Verify are the guard; mutation, edge and oracle scores show it

## PLAN
strategy: red tests for load, resolution, argv and records; add the fields and plumb `model` through pilot → run_reps → run_pilot → execute_wm → build_argv; write the two arm tomls and the variant skill; then round 6
check: python3 -m pytest -q benchmark/tests/test_arm_models.py benchmark/tests/test_arms.py
regression: python3 -m pytest -q benchmark/tests

## CHECKS
- C1 covers: M1 · acceptance · benchmark/tests/test_arm_models.py::test_arm_toml_model_and_advisor_default_empty · falsifier: the loader ignores the new keys
- C2 covers: M2 R:DEFAULT · acceptance · benchmark/tests/test_arm_models.py::test_resolved_model_prefers_arm_then_flag_then_pin · falsifier: `--model` overrides the advisor arm's own Haiku
- C3 covers: M3 · acceptance · benchmark/tests/test_arm_models.py::test_argv_carries_model_effort_and_advisor_only_when_set · falsifier: `--advisor` with an empty value on every arm
- C4 covers: M4 M2 · acceptance · benchmark/tests/test_arm_models.py::test_execute_wm_passes_and_records_the_resolved_model · falsifier: argv uses the new model but the record still says claude-sonnet-5
- C5 covers: M2 · acceptance · benchmark/tests/test_arm_models.py::test_run_all_cli_accepts_model · falsifier: the flag parses but never reaches run_reps
- C6 covers: M5 R:ARMCOUNT · acceptance · benchmark/tests/test_arm_models.py::test_lean_and_advisor_arms_load · falsifier: the advisor arm runs the shipped skill, or the lean arm pins a model
- C7 covers: M6 · acceptance · benchmark/tests/test_arm_models.py::test_lean_variant_changes_exactly_three_things · falsifier: a variant that also drops a check rule or the seal
- C9 covers: R:NO_LIVE · acceptance · benchmark/tests/test_arm_models.py::test_no_test_can_launch_the_real_claude · falsifier: a guard that only wraps build_argv, which a test calling the launcher directly walks past
- C8 covers: R:ARMCOUNT · regression · benchmark/tests/test_arms.py::test_all_arms_validate_with_fairness_parity · falsifier: arms added with diverging fairness fields

## LOG
- 2026-09-30 refreeze in build: the autouse guard in benchmark/tests/conftest.py wraps `build_argv` with a two-argument signature and raises whenever no agent is injected, so C4 — which replaces `_invoke_once` and launches nothing — cannot run. The guard moves to the launch layer (refuse a process whose binary is `claude`), keeping its purpose; scope widens to conftest.py; R:NO_LIVE and C9 make the safety property a sealed check.
- 2026-09-30 refreeze in build: C6 demanded the arm's first steps equal all four add-4 steps, which puts the variant copy after `workspace_git.py`'s baseline commit — the agent would then see a modified SKILL.md in its working tree, fouling Verify's clean-tree run. C6 now requires every add-4 step in order, with the copy between the install and the baseline commit; M5 says so.
- 2026-09-30 refreeze in build: two older tests pin the old two-argument `build_argv` shape — test_session_mode's spy takes (prompt, agent_cmd), and test_wv2_family greps core.py for the literal `"model": PINNED_MODEL`. Their intent (every WM starts a fresh conversation; every record stamps the model it ran on) is unchanged: the spy passes new arguments through and the grep looks for the resolved model. Scope widens to both files.

## EVIDENCE
verdict: PASS
- seal: since the last refreeze (d41beda0) no sealed check file changed; `benchmark/tests/conftest.py` changed in the build commit exactly as refreeze b72e7ed2 set out (the guard wraps `subprocess.Popen` and refuses a `claude` binary), and C9 holds it
- fresh: `python3 -m pytest -q benchmark/tests/test_arm_models.py benchmark/tests/test_arms.py` → 10 passed (C1–C9)
- regression: `python3 -m pytest -q benchmark/tests` → 530 passed, 12 skipped
- probe (M4, live): round-6 records stamp the model each run used — add-4-advisor wm1/amb1 `claude-haiku-4-5-20251001` + advisor `claude-sonnet-5-5`; add-4-lean `claude-sonnet-5-5`, no advisor
- probe (M5, live): the add-4-advisor wm1 workspace's SKILL.md carries the variant (`model: haiku` ×1); in add-4-lean wm1 the SKILL.md's only commit is `chore(bench): workspace baseline` and `git status .claude` is clean — the variant landed before the baseline
- residue: measurement results, not harness defects — the Haiku main session never consulted its advisor in the first two runs, although a forced consult does bill Sonnet 5.5 ($0.15), and the lean arm's Build did not hand off to a Haiku subagent; both go to the round-6 report
45 changes: 45 additions & 0 deletions .add/tasks/bench-sonnet-only.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
---
type: Task
title: ADD's benchmark arms run on Sonnet only — the Haiku advisor arm retires and the lean skill keeps Build in the Sonnet session
status: done
kind: change
risks: [measurement-validity]
scope: [benchmark/arms/, benchmark/tests/test_arm_models.py, benchmark/tests/test_arms.py]
gives: [the lean variant as a two-edit Sonnet-only skill]
---
## CARD
goal: every ADD arm runs each beat on one Sonnet model; the lean variant keeps only the two edits that cut work
why: round 6 (Sonnet 5.5, 3 reps) — the Haiku main session never consulted its Sonnet advisor in 6 runs, failed all 3 wm1 apps on Flask/FastAPI, and cost 1.6–1.8× plain add-4 at twice the wall time; the lean arm never handed Build to a Haiku subagent (0 of 6), so that line is dead text

## RULES
- M1 the `add-4-advisor` arm retires: it leaves `ARM_NAMES` and its toml is deleted (from: request — "use sonnet only")
- M2 the lean variant differs from the shipped skill in exactly two places — stub only what the checks import; pick the lead persona by grepping `.add/personas/` — and names no Haiku model; Build and Verify stay in the main session (from: request; PILOT r6 — 0 subagent hand-offs in 6 lean runs)
- M3 an arm that pins a model pins a Sonnet one (from: request — "use sonnet only")
- R:KEEP the `model` / `advisor` arm keys and `run-all --model` stay — `--model claude-sonnet-5-5` needs them (from: .add/tasks/bench-model-switch.md M1–M4)
- R:ARMCOUNT `ARM_NAMES` shrinks from 10 to 9; the fairness fields stay identical across all arms (from: benchmark/tests/test_arms.py::test_all_arms_validate_with_fairness_parity)

## ASSUMPTIONS
- A1 [which] "sonnet only" means the arms and the ADD skill, not the harness plumbing → R:KEEP; the advisor key stays generic and unused
- A2 [experience] round-6 lean data still describes the two-edit variant → the Haiku line was never acted on in 6 runs, so the model saw it but did nothing with it; the report says so

## PLAN
strategy: re-aim C6/C7 of bench-model-switch and add an all-arms Sonnet check (red); delete the advisor toml, drop it from ARM_NAMES, cut the variant's Haiku hand-off back to the shipped Build line
check: python3 -m pytest -q benchmark/tests/test_arm_models.py benchmark/tests/test_arms.py
regression: python3 -m pytest -q benchmark/tests

## CHECKS
- C1 covers: M1 R:ARMCOUNT · acceptance · benchmark/tests/test_arm_models.py::test_lean_arm_loads_and_advisor_arm_retired · falsifier: the advisor toml deleted but its name still in ARM_NAMES
- C2 covers: M2 · acceptance · benchmark/tests/test_arm_models.py::test_lean_variant_changes_exactly_two_things · falsifier: the Haiku hand-off line survives
- C3 covers: M3 · acceptance · benchmark/tests/test_arm_models.py::test_every_pinned_arm_model_is_sonnet · falsifier: a new arm pins claude-haiku-4-5
- C4 covers: R:KEEP · regression · benchmark/tests/test_arm_models.py::test_argv_carries_model_effort_and_advisor_only_when_set · falsifier: the advisor plumbing removed with the arm
- C5 covers: R:ARMCOUNT · regression · benchmark/tests/test_arms.py::test_all_arms_validate_with_fairness_parity · falsifier: the count left at 10

## LOG

## EVIDENCE
verdict: PASS
- seal: `git diff e11b2a5a HEAD -- benchmark/tests .add/tasks` is empty before this record
- fresh: `python3 -m pytest -q benchmark/tests/test_arm_models.py benchmark/tests/test_arms.py` → 11 passed (C1–C5; red at freeze: C1 C2 C3 C5 failed, each for its falsifier's reason)
- regression: `python3 -m pytest -q benchmark/tests` → 531 passed, 12 skipped
- probe (M2): `diff add-method/skill/add/SKILL.md benchmark/arms/variants/add-4-lean/SKILL.md` → 2 hunks (stubs, persona grep); no "haiku" in the variant
- residue: `add-4-advisor` survives only in the two task records and the test asserting it is gone; round-6 run records keep it as history
119 changes: 119 additions & 0 deletions benchmark/PILOT-4v3-2026-09-30-r6.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,119 @@
# Pilot round 6 — Sonnet 5.5, and does switching models by beat cut Direction? (2026-09-30)

Question: Direction took 56% of ADD's wall time (4.1 of 7.3 min) in round 5, on `claude-sonnet-5`.
Does a newer model, a leaner skill, or a cheaper main model with a stronger advisor cut its time
and cost without losing quality?

Arms, all at harness 31b6f1aa, `--effort medium`, wm1 + amb1, **n = 3 per arm per workload. Not
significant.**

- **vanilla** and **add-4** (ADD 4.0 as shipped) on `claude-sonnet-5-5`.
- **add-4-lean** on `claude-sonnet-5-5`, running a variant of the skill with three edits:
- stub only what the checks import;
- pick the lead persona by grepping `.add/personas/`, not by reading the 65 KB index;
- hand Build to one Haiku subagent.
- **add-4-advisor**: the same variant on a `claude-haiku-4-5` main session, with `--advisor
claude-sonnet-5-5`.

Runs are in `benchmark/runs-4v3-2026-09-30-r6/` (gitignored). Both arms load the operator's
`~/.claude`, as in rounds 4–5.

## Results

| wm1 | vanilla | add-4 | add-4-lean | add-4-advisor (Haiku) |
|---|---|---|---|---|
| oracle pass rate | 1.00 · 1.00 · 1.00 | 1.00 · 1.00 · 1.00 | 1.00 · 1.00 · 1.00 | **0 · 0 · 0** |
| cost | $0.35 · $0.30 · $0.28 ($0.31) | $0.54 · $0.54 · $0.48 ($0.52, 1.7×) | $0.53 · $0.44 · $0.48 ($0.48) | $0.66 · $1.19 · $0.91 ($0.92) |
| wall time | 0.8 min | 4.1 min | 2.5 min | 6.8 min |
| Direction (wall · API messages) | — | 1.1 min · 3 | 1.0 min · 3 | 2.7 min · 24 |
| edge robustness (held out, /19) | 19 · 19 · 19 | 19 · 19 · 19 | 19 · 19 · 19 | 0 (app never started) |
| mutation score (own tests) | 0.92 · 0.83 · 0.75 (0.83) | 0.75 · 0.58 · 0.92 (0.75) | 0.88 · 0.75 · 0.58 (0.74) | 0.62 × 3 |
| own tests | 12.7 per run | 20.7 | 24.3 | 43.3 |

| amb1 | vanilla | add-4 | add-4-lean | add-4-advisor (Haiku) |
|---|---|---|---|---|
| oracle pass rate | 0.88 · 1.00 · 1.00 | 1.00 · 1.00 · 1.00 | 1.00 · 1.00 · 1.00 | 0.88 · 1.00 · 0.75 |
| cost | $0.37 · $0.28 · $0.34 ($0.33) | $0.78 · $0.63 · $0.61 ($0.68, 2.1×) | $0.60 · $0.65 · $0.72 ($0.66) | $0.71 · $1.32 · $1.26 ($1.10) |
| wall time | 0.85 min | 3.4 min | 2.9 min | 8.4 min |
| Direction (wall · API messages) | — | 1.7 min · 4 | 1.7 min · 5 | 3.4 min · 29 |
| planted ambiguities right, of 7 | 4 · 5 · 4 | 5 · 6 · 6 | 6 · 5 · 6 | 4 · 5 · 3 |
| edge robustness (held out, /14) | 14 · 14 · 14 | 14 · 14 · 14 | 14 · 14 · 14 | 10 · 10 · 10 |
| mutation score (own tests) | 0.83 · 0.79 · 0.67 (0.76) | 0.75 · 0.75 · 0.88 (0.79) | 0.79 · 0.96 · 0.75 (0.83) | 0.38 · 0.58 · 0.54 (0.50) |

Every ADD run's evidence claims matched a rerun (e.g. add-4 wm1 41→41, 66→66, 62→62). Two lean
amb1 runs and one advisor run wrote no parseable count.

## Direction: time and cost

**The model upgrade did most of the work.** The same ADD 4.0 skill went from Sonnet 5 (round 5)
to Sonnet 5.5:

- Direction: 4.1 min (round-5 mean) → 1.1 min on wm1 and 1.7 min on amb1.
- API messages per Direction: 23 → 3. Sonnet 5.5 packs a beat's tool calls into far fewer messages.
- Cost: $1.67 → $0.52 on wm1, and $1.31 → $0.68 on amb1.
- ADD relative to vanilla: 2.8× → 1.7× on wm1, and 2.9× → 2.1× on amb1.

On Sonnet 5.5 no run opened the 65 KB persona index. Direction wrote 0–3 files instead of about 11.

**Where ADD's bill goes now (add-4, wm1, $0.52 a run):**

- **Fixed context.** The first message writes the ~35k-token system context (Claude Code, the
operator's `~/.claude`, the skill) to the cache: **$0.14, over a quarter of the bill**. Vanilla
pays the same $0.17 for its ~41k-token context. This is harness and operator setup, not ADD.
- **Output.** 17k output tokens against vanilla's 7k (+$0.10): the task file, 21 tests against 13,
and the evidence.
- **Cache reads.** 529k against 213k (+$0.06): more turns re-read the context.

By transcript pricing Direction takes 60–67% of ADD's dollars, but most of that is the fixed first
message landing in its beat. Beyond it, Direction costs about $0.08–0.13 a run.

**Where the time goes now.** On wm1, Build is the long beat (2.75 of 4.1 min). The lean skill cut
Build to 1.3 min; Direction did not move (1.0 against 1.1 min). With fewer stubs to replace, Build
writes the code once.

## The model switch — refuted

- **Haiku never consulted its advisor.** In 6 runs, 0 advisor calls and $0 billed to Sonnet. A
probe that told Haiku to consult did bill Sonnet, **$0.15 for one consult**, over a quarter of a
whole Sonnet 5.5 ADD run.
- **Haiku costs more, not less.** Direction took 24–29 messages against 3–5. Runs cost $0.92–1.10
against $0.52–0.68, at about twice the wall time.
- **Haiku broke the entry contract.** All 3 wm1 apps were built on Flask or FastAPI in a private
venv, and none starts under the oracle's stdlib-only `python -S`. Sonnet 5.5 met the contract in
9 of 9 wm1 runs.
- **The Build hand-off to a Haiku subagent never happened.** 0 of 6 lean runs, which fits rounds
4–5: a rule changes behaviour when it is an artifact the model must write, and a line of prose
about how to delegate does not.

The advisor arm was retired after this round, and the lean variant's Build line went back to the
shipped text (task `bench-sonnet-only`). The Haiku line was never acted on, so the lean numbers
above describe the two remaining edits.

## What this does not show

On Sonnet 5.5 the two workloads **saturate**:

- Every Sonnet arm reached 19/19 and 14/14 on the held-out edges.
- Vanilla wrote tests in 6 of 6 runs; round 5 had two vanilla runs with no tests at all.
- Mutation scores sit within noise.

ADD's measured lead on this model narrows to:

- **The ambiguous spec:** 5.7 and 5.7 of 7 planted ambiguities handled right, against 4.3 for vanilla.
- **Evidence claims:** every count matched a rerun.

Separating the arms on quality at this model tier needs harder workloads (the hv hard-cases track).

## Recommendations

1. **Run ADD on Sonnet 5.5, one model for every beat.** It is the largest measured cut in
Direction's time and cost, and the only one that needed no change to the method.
2. **Adopt the lean skill's two edits in `add-method/skill/add/SKILL.md`** as their own task:
- stub only what the checks import;
- pick the persona by grep.

At n = 3 they cut wall time 14–38% and cost 3–8%, with no quality loss on either workload.
3. **Don't switch models by beat, and don't pair a smaller main model with an advisor.** Both were
measured here and both lost.
4. **The biggest remaining fixed cost is the operator's `~/.claude` context** (about $0.14–0.17 a
run on either arm). A clean `HOME` for benchmark runs would measure the method without it.
Loading
Loading