Skip to content

ADD at low effort vs vanilla: isolated benchmark, SWE-bench Lite pilot, Quick lane for bounded fixes - #232

Open
TinDang97 wants to merge 18 commits into
mainfrom
feat/add-low-effort-bench
Open

TinDang97 wants to merge 18 commits into
mainfrom
feat/add-low-effort-bench

Conversation

@TinDang97

@TinDang97 TinDang97 commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

ADD at --effort low against vanilla Claude Code at medium, both on Sonnet 5.5, measured with the operator's config isolated. Acting on that measurement, bounded fixes now take the Quick lane.

  • Benchmark harness
    • Per-arm effort, recorded in every run record.
    • --setting-sources project,local, so the operator's ~/.claude stays out of measured sessions.
    • The SWE-bench Lite runner ported to ADD 4.0, with a fixed --sample slice, --workers, kept transcripts, and patches that include new files.
  • Finding: rounds 4–7 measured the operator, not ADD.
    • The security-guidance plugin ran an LLM review on every git commit, 100–185 s each. Only ADD commits, so only ADD paid it.
    • The user CLAUDE.md ("interview me") made vanilla halt on the ambiguous spec.
  • Skill (task lean-bounded-fixes, verify RISK-ACCEPTED)
    • The floor is a change of shape to a consumed surface, not a touch, so bounded fixes take the Quick lane.
    • ASSUMPTIONS hold only real silences.
    • The persona-index path now resolves.
    • A contradiction rule was measured harmful and reverted.
  • READMEs: the isolated price (1.3–1.9× the dollars, 1.8–2.4× the minutes), a recommendation of --effort low, and the pooled mutation gap.

Results

vanilla · medium ADD · low
wm1 · amb1 oracle / held-out edges 1.00 · 22/22 1.00 · 22/22
seeded bugs caught by own tests (rounds 8–9 pooled) 0.78 · 0.43 0.86 · 0.81
$ per run (wm1 · amb1) $0.21 · $0.17 $0.28 · $0.32
SWE-bench Lite, 30-instance slice resolved 21/30 ($0.13 per resolved) 23/30 (4.0.0 skill, $0.33) → 21/30 (lean skill, $0.24)

Round 10: the 2×2 effort sweep (final value-final skill, verify PASS)

vanilla · low vanilla · medium ADD · low ADD · medium
SWE Lite resolved, of 30 21 20 24 22
$ per resolved issue $0.12 $0.14 $0.23 $0.32
s per issue 20 27 59 80
ran repo tests / shipped a test 5 / 1 11 / 1 30 / 30 30 / 30
amb1 ambiguities, of 7 4.3 4.3 5.3 5.7
mutation (wm1 · amb1) 0.89 · 0.36 0.75 · 0.67 0.92 · 0.78 0.75 · 0.75

ADD-low resolves every issue vanilla-medium resolves, plus 4 more. Medium effort bought neither tool anything. Page: benchmark/results/2026-10-add-4.0-low-effort-vs-vanilla.md.

All n = 3 or n = 30: directional, not significant. Details:

  • benchmark/PILOT-4v3-2026-10-02-r7.md (rounds 7–8)
  • benchmark/SWE-LITE-PILOT-2026-10-03.md
  • benchmark/PROMPT-AUDIT-2026-10-02.md
  • benchmark/results/2026-10-add-4.0-low-effort-vs-vanilla.md

Open risk

The Quick lane cut SWE cost 35%, but resolved went 23 → 21 of 30, within noise. A full 300-instance run (~$70) settles it.

Test plan

  • cd add-method && python3 -m pytest -q: 159 pass
  • python3 -m pytest -q benchmark/tests: 547 pass, 12 skipped
  • The three skill trees are byte-identical; SKILL.md is 199 of 200 lines
  • A fresh pilotspace-add init resolves the persona path the skill names

The meter hard-coded `--effort medium` for every arm, so one question could
not be asked: does ADD's ceremony at low effort hold the quality raw Claude
Code needs medium for, on the same model? Effort now resolves like model —
the arm's `effort`, else `run-all --effort`, else medium — is validated at
load time (low|medium|high|xhigh|max) and stamped into artifacts.effort.

Two new arms: add-4-low (shipped ADD 4.0 at low) and add-4-audited-low
(the prompt-audited SKILL.md variant at low: numeric turn counts out, stub
only what the checks import, personas picked by grep at their installed
`.add/` path). Control arms keep the medium default.

Tests: benchmark/tests/test_arm_effort.py (red first, 9 cases); suite 540 pass.

author: Tin Dang
Round 7 put ~175 s of each ADD run's ~200 s of test-command wall time
outside pytest itself (pytest reported 19-31 s per run): Verify started the
app with `python3 -m app &` to probe it by hand, and an unkilled or crashed
server held the Bash call open until it was stopped or timed out at 120 s.

add-4-probe-low = the audited variant plus one Verify line: probes are test
files in `.add/tasks/<slug>.probes/` run by the test command; an app the
agent must start logs to a file and is killed in the same command. Low effort.

Tests: test_arm_effort.py +1 (red first); suite 541 pass.

author: Tin Dang
…sions

A live stall watch on round 8 caught the operator's `security-guidance`
plugin: its PostToolUse hook runs an Opus security review on every
`git commit`, holding the agent's Bash call for 100-185 s and billing
outside the run record. ADD commits 2-4 times per run (freeze, build,
verify); vanilla never commits — so the operator's plugin was measured as
ADD's wall time in every round since 4.0. The user CLAUDE.md also reached
every arm ("interview me until 95% confidence" in a headless run).

`--setting-sources project,local` drops user plugins, hooks and
~/.claude/CLAUDE.md (verified: plugins 17 -> 4 org-managed, context
35k -> 24k tokens, the user rule no longer quotable). Organisation-managed
plugins still load, equally for every arm.

Tests: test_arm_effort.py +1 (red first); suite 542 pass.

author: Tin Dang
The runner still installed ADD 2.0 (`.add/tooling/add.py`, `--force`), drove
the retired engine, pinned `--effort medium` for both arms and loaded the
operator's user settings. Now:

- add arm: the ADD 4.0 installer only; the prompt drives
  `.claude/skills/add/SKILL.md` (freeze/verify commits, a reproducing check,
  host tests as the regression floor, headless — no human)
- ARM_EFFORT: add=low, vanilla=medium on claude-sonnet-5-5, via the wm
  bench's own argv builder (so --setting-sources project,local applies)
- collect_patch: `git add -A` + `diff --cached <base>`, so a NEW file the
  fix creates is no longer dropped from the prediction
- --sample N --seed S: a fixed slice of all 300 Lite ids; --workers N
- each instance keeps transcript.jsonl (the trajectory a submission needs)
- evaluation documented on Modal (x86) instead of local docker

Tests: test_swe_smoke.py rewritten for 4.0 (red first, 7 new failing);
suite 544 pass.

author: Tin Dang
The per-id datasets-server /filter endpoint returned HTTP 500 through every
retry; the paged /rows fetch that lists the 300 ids already carries each
row, so it now fills instances.json and /filter is never reached.

Tests: test_swe_smoke.py +2 (red first); suite pass.

author: Tin Dang
…ator-config finding

Round 8 (isolated, sequential) supersedes round 7: the operator's
security-guidance plugin reviewed every git commit (ADD-only cost) and the
user CLAUDE.md made vanilla halt on amb1. Verdicts: ADD at low holds and
pays on test strength (amb1 mutation 0.83 vs 0.19) at 1.3-1.9x cost; low
effort is the cheaper default; ship only the audit's path fix; the probe
rule is not needed.

author: Tin Dang
…vanilla-medium 21/30

ADD resolves a strict superset (+2: scikit-learn-14087, sympy-14817) at
2.7x cost ($0.33 vs $0.13 per resolved); McNemar p=0.5, directional.
Scored with the official harness locally (swebench 5.0.2, x86 images
under Rosetta) since every swebench release's Modal path is broken.

author: Tin Dang
…s only where the spec is silent, contradictions in the caller's favour, honest README price

author: Tin Dang
…nd the benchmark variant tests

Both pinned their checks to artifacts this task moves on purpose (the
2026-09 results page; the live SKILL.md). They now read the cited results
page and a 4.0.0 base snapshot — same assertions, stable sources.

author: Tin Dang
… silences, contradictions in the caller's favour

- floor = a change of shape to a consumed surface, not a touch; a fix that
  restores intended behavior without changing its shape is not floor work
- a Quick commit body records lane + red->green evidence (an artifact)
- ASSUMPTIONS: settled dimensions get no line, '- none — <why>' is valid;
  contradictory requirements -> the caller-in-control reading, first
- personas index path is the installed .add/personas-index/use-when.md
- READMEs quote round 8's isolated price, recommend --effort low, name the
  security-guidance contamination; new results page backs every number
- SWE runner: the ADD prompt lets the skill size the work
- compensations to stay at 200 lines: two persona/report glosses dropped
  (still in references/personas.md), 'performance' out of the risk
  examples list
- guard tests read the cited results page / a 4.0.0 base snapshot

Tests: add-method 160 pass, benchmark 547 pass.

author: Tin Dang
…utation claim

Round 9: the caller-in-control reading ('reject') left amb1's waitlist
requirements 4-6 with nothing to act on — ambiguities handled 5.3 -> 4.0,
one oracle 0.88. Round 9 vanilla scored 0.67 on amb1 mutation, so round
8's 0.83 vs 0.19 overstated the gap; the READMEs quote rounds 8-9 pooled.

author: Tin Dang
…utation gap

Round 9 measured the caller-in-control reading as net harmful on amb1
(ambiguities 5.3 -> 4.0, one oracle 0.88). The READMEs and results page
now quote rounds 8-9 pooled (amb1 0.81 vs 0.43), disclosing that round 9's
ADD arm ran the reverted draft.

Tests: add-method 159 pass, benchmark 547 pass.

author: Tin Dang
Seal intact since refreeze 0a7fa83; check 5+22 pass, regression 159+547
pass, three refute probes pass on a fresh install. Quick lane on bounded
fixes: SWE Lite cost -35% (2.7x -> 1.8x vanilla), 30/30 patches with
tests; resolved 23 -> 21 of 30 (within noise). Open risk owned by Tin Dang:
settle with the full 300-instance run. The commit-body artifact (M2) did not
transfer; the contradiction rule (M4) was reverted as harmful.

author: Tin Dang
Comment thread benchmark/runner/core.py
from benchmark.check_isolation import find_leaks
from benchmark.check_isolation import main as check_isolation_main
from benchmark.runner.agent import PINNED_MODEL, build_argv, resolve_model
from benchmark.runner.agent import PINNED_MODEL, build_argv, resolve_effort, resolve_model
…d found: step; refute for security work

Yield audit of rounds 8-9 (24 task files) and the SWE-bench pilots:
falsifier-per-check followed 24/24 and carries the mutation gain; found:
0/24; refute probes 1/24 with 24/24 PASS; Quick tests 1.3 asserts/fix and
the 3 lost SWE issues were wrong fixes. READMEs: a measured row for tests
run (30/30 vs 7/30) and shipped (30/30 vs 2/30).

Tests: add-method 163 pass, benchmark 547 pass.

author: Tin Dang
Round 10 runs the final value-final skill and raw Claude Code at --effort low and
medium on claude-sonnet-5-5, isolated with --setting-sources project,local. wm1
and amb1 use n = 3 per cell. SWE-bench Lite uses the same 30 instances (seed 0),
scored locally by the official harness (swebench 5.0.2).

SWE resolved: vanilla-low 21, vanilla-medium 20, ADD-low 24, ADD-medium 22.
ADD-low's set contains every vanilla-medium resolve plus 4 more. Cost per resolved
issue is $0.23 for ADD-low against $0.14 for vanilla-medium. Medium effort
bought neither tool anything.

The results page, the SWE pilot doc, both READMEs and the CHANGELOG now carry
the grid.

author: Tin Dang
Seal intact since freeze 6fc12fe. Check: 4 tests pass. Regression: 163 + 549
pass. Round 10 re-measure, on the same 30 SWE-bench Lite instances: ADD-low
resolves 24/30, a superset of vanilla-medium's 20, at $0.23 per resolved issue.
That recovers the lean skill's 21 at a lower cost than 4.0.0. On wm1 and amb1,
ADD-low leads on ambiguities (5.3 vs 4.3) and on mutation, and correctness ties.

author: Tin Dang

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant