ADD at low effort vs vanilla: isolated benchmark, SWE-bench Lite pilot, Quick lane for bounded fixes - #232
Open
TinDang97 wants to merge 18 commits into
Open
ADD at low effort vs vanilla: isolated benchmark, SWE-bench Lite pilot, Quick lane for bounded fixes#232TinDang97 wants to merge 18 commits into
TinDang97 wants to merge 18 commits into
Conversation
The meter hard-coded `--effort medium` for every arm, so one question could not be asked: does ADD's ceremony at low effort hold the quality raw Claude Code needs medium for, on the same model? Effort now resolves like model — the arm's `effort`, else `run-all --effort`, else medium — is validated at load time (low|medium|high|xhigh|max) and stamped into artifacts.effort. Two new arms: add-4-low (shipped ADD 4.0 at low) and add-4-audited-low (the prompt-audited SKILL.md variant at low: numeric turn counts out, stub only what the checks import, personas picked by grep at their installed `.add/` path). Control arms keep the medium default. Tests: benchmark/tests/test_arm_effort.py (red first, 9 cases); suite 540 pass. author: Tin Dang
Round 7 put ~175 s of each ADD run's ~200 s of test-command wall time outside pytest itself (pytest reported 19-31 s per run): Verify started the app with `python3 -m app &` to probe it by hand, and an unkilled or crashed server held the Bash call open until it was stopped or timed out at 120 s. add-4-probe-low = the audited variant plus one Verify line: probes are test files in `.add/tasks/<slug>.probes/` run by the test command; an app the agent must start logs to a file and is killed in the same command. Low effort. Tests: test_arm_effort.py +1 (red first); suite 541 pass. author: Tin Dang
…sions
A live stall watch on round 8 caught the operator's `security-guidance`
plugin: its PostToolUse hook runs an Opus security review on every
`git commit`, holding the agent's Bash call for 100-185 s and billing
outside the run record. ADD commits 2-4 times per run (freeze, build,
verify); vanilla never commits — so the operator's plugin was measured as
ADD's wall time in every round since 4.0. The user CLAUDE.md also reached
every arm ("interview me until 95% confidence" in a headless run).
`--setting-sources project,local` drops user plugins, hooks and
~/.claude/CLAUDE.md (verified: plugins 17 -> 4 org-managed, context
35k -> 24k tokens, the user rule no longer quotable). Organisation-managed
plugins still load, equally for every arm.
Tests: test_arm_effort.py +1 (red first); suite 542 pass.
author: Tin Dang
The runner still installed ADD 2.0 (`.add/tooling/add.py`, `--force`), drove the retired engine, pinned `--effort medium` for both arms and loaded the operator's user settings. Now: - add arm: the ADD 4.0 installer only; the prompt drives `.claude/skills/add/SKILL.md` (freeze/verify commits, a reproducing check, host tests as the regression floor, headless — no human) - ARM_EFFORT: add=low, vanilla=medium on claude-sonnet-5-5, via the wm bench's own argv builder (so --setting-sources project,local applies) - collect_patch: `git add -A` + `diff --cached <base>`, so a NEW file the fix creates is no longer dropped from the prediction - --sample N --seed S: a fixed slice of all 300 Lite ids; --workers N - each instance keeps transcript.jsonl (the trajectory a submission needs) - evaluation documented on Modal (x86) instead of local docker Tests: test_swe_smoke.py rewritten for 4.0 (red first, 7 new failing); suite 544 pass. author: Tin Dang
The per-id datasets-server /filter endpoint returned HTTP 500 through every retry; the paged /rows fetch that lists the 300 ids already carries each row, so it now fills instances.json and /filter is never reached. Tests: test_swe_smoke.py +2 (red first); suite pass. author: Tin Dang
…ator-config finding Round 8 (isolated, sequential) supersedes round 7: the operator's security-guidance plugin reviewed every git commit (ADD-only cost) and the user CLAUDE.md made vanilla halt on amb1. Verdicts: ADD at low holds and pays on test strength (amb1 mutation 0.83 vs 0.19) at 1.3-1.9x cost; low effort is the cheaper default; ship only the audit's path fix; the probe rule is not needed. author: Tin Dang
…vanilla-medium 21/30 ADD resolves a strict superset (+2: scikit-learn-14087, sympy-14817) at 2.7x cost ($0.33 vs $0.13 per resolved); McNemar p=0.5, directional. Scored with the official harness locally (swebench 5.0.2, x86 images under Rosetta) since every swebench release's Modal path is broken. author: Tin Dang
…s only where the spec is silent, contradictions in the caller's favour, honest README price author: Tin Dang
…nd the benchmark variant tests Both pinned their checks to artifacts this task moves on purpose (the 2026-09 results page; the live SKILL.md). They now read the cited results page and a 4.0.0 base snapshot — same assertions, stable sources. author: Tin Dang
… silences, contradictions in the caller's favour - floor = a change of shape to a consumed surface, not a touch; a fix that restores intended behavior without changing its shape is not floor work - a Quick commit body records lane + red->green evidence (an artifact) - ASSUMPTIONS: settled dimensions get no line, '- none — <why>' is valid; contradictory requirements -> the caller-in-control reading, first - personas index path is the installed .add/personas-index/use-when.md - READMEs quote round 8's isolated price, recommend --effort low, name the security-guidance contamination; new results page backs every number - SWE runner: the ADD prompt lets the skill size the work - compensations to stay at 200 lines: two persona/report glosses dropped (still in references/personas.md), 'performance' out of the risk examples list - guard tests read the cited results page / a 4.0.0 base snapshot Tests: add-method 160 pass, benchmark 547 pass. author: Tin Dang
…utation claim
Round 9: the caller-in-control reading ('reject') left amb1's waitlist
requirements 4-6 with nothing to act on — ambiguities handled 5.3 -> 4.0,
one oracle 0.88. Round 9 vanilla scored 0.67 on amb1 mutation, so round
8's 0.83 vs 0.19 overstated the gap; the READMEs quote rounds 8-9 pooled.
author: Tin Dang
…utation gap Round 9 measured the caller-in-control reading as net harmful on amb1 (ambiguities 5.3 -> 4.0, one oracle 0.88). The READMEs and results page now quote rounds 8-9 pooled (amb1 0.81 vs 0.43), disclosing that round 9's ADD arm ran the reverted draft. Tests: add-method 159 pass, benchmark 547 pass. author: Tin Dang
Seal intact since refreeze 0a7fa83; check 5+22 pass, regression 159+547 pass, three refute probes pass on a fresh install. Quick lane on bounded fixes: SWE Lite cost -35% (2.7x -> 1.8x vanilla), 30/30 patches with tests; resolved 23 -> 21 of 30 (within noise). Open risk owned by Tin Dang: settle with the full 300-instance run. The commit-body artifact (M2) did not transfer; the contradiction rule (M4) was reverted as harmful. author: Tin Dang
| from benchmark.check_isolation import find_leaks | ||
| from benchmark.check_isolation import main as check_isolation_main | ||
| from benchmark.runner.agent import PINNED_MODEL, build_argv, resolve_model | ||
| from benchmark.runner.agent import PINNED_MODEL, build_argv, resolve_effort, resolve_model |
… the model never does author: Tin Dang
…d found: step; refute for security work Yield audit of rounds 8-9 (24 task files) and the SWE-bench pilots: falsifier-per-check followed 24/24 and carries the mutation gain; found: 0/24; refute probes 1/24 with 24/24 PASS; Quick tests 1.3 asserts/fix and the 3 lost SWE issues were wrong fixes. READMEs: a measured row for tests run (30/30 vs 7/30) and shipped (30/30 vs 2/30). Tests: add-method 163 pass, benchmark 547 pass. author: Tin Dang
…2x2 effort sweep author: Tin Dang
Round 10 runs the final value-final skill and raw Claude Code at --effort low and medium on claude-sonnet-5-5, isolated with --setting-sources project,local. wm1 and amb1 use n = 3 per cell. SWE-bench Lite uses the same 30 instances (seed 0), scored locally by the official harness (swebench 5.0.2). SWE resolved: vanilla-low 21, vanilla-medium 20, ADD-low 24, ADD-medium 22. ADD-low's set contains every vanilla-medium resolve plus 4 more. Cost per resolved issue is $0.23 for ADD-low against $0.14 for vanilla-medium. Medium effort bought neither tool anything. The results page, the SWE pilot doc, both READMEs and the CHANGELOG now carry the grid. author: Tin Dang
Seal intact since freeze 6fc12fe. Check: 4 tests pass. Regression: 163 + 549 pass. Round 10 re-measure, on the same 30 SWE-bench Lite instances: ADD-low resolves 24/30, a superset of vanilla-medium's 20, at $0.23 per resolved issue. That recovers the lean skill's 21 at a lower cost than 4.0.0. On wm1 and amb1, ADD-low leads on ambiguities (5.3 vs 4.3) and on mutation, and correctness ties. author: Tin Dang
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ADD at
--effort lowagainst vanilla Claude Code at medium, both on Sonnet 5.5, measured with the operator's config isolated. Acting on that measurement, bounded fixes now take the Quick lane.effort, recorded in every run record.--setting-sources project,local, so the operator's~/.claudestays out of measured sessions.--sampleslice,--workers, kept transcripts, and patches that include new files.security-guidanceplugin ran an LLM review on everygit commit, 100–185 s each. Only ADD commits, so only ADD paid it.CLAUDE.md("interview me") made vanilla halt on the ambiguous spec.lean-bounded-fixes, verify RISK-ACCEPTED)--effort low, and the pooled mutation gap.Results
Round 10: the 2×2 effort sweep (final
value-finalskill, verify PASS)ADD-low resolves every issue vanilla-medium resolves, plus 4 more. Medium effort bought neither tool anything. Page:
benchmark/results/2026-10-add-4.0-low-effort-vs-vanilla.md.All n = 3 or n = 30: directional, not significant. Details:
benchmark/PILOT-4v3-2026-10-02-r7.md(rounds 7–8)benchmark/SWE-LITE-PILOT-2026-10-03.mdbenchmark/PROMPT-AUDIT-2026-10-02.mdbenchmark/results/2026-10-add-4.0-low-effort-vs-vanilla.mdOpen risk
The Quick lane cut SWE cost 35%, but resolved went 23 → 21 of 30, within noise. A full 300-instance run (~$70) settles it.
Test plan
cd add-method && python3 -m pytest -q: 159 passpython3 -m pytest -q benchmark/tests: 547 pass, 12 skippedSKILL.mdis 199 of 200 linespilotspace-add initresolves the persona path the skill names