Skip to content

docs: the book matches 4.0, and the READMEs trade ADD against vanilla Claude with an animated side-by-side flow - #231

Merged
TinDang97 merged 9 commits into
mainfrom
docs/add-4-value
Sep 30, 2026
Merged

TinDang97 merged 9 commits into
mainfrom
docs/add-4-value

Conversation

@TinDang97

@TinDang97 TinDang97 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The book and READMEs now match the merged ADD 4.0 skill. Every 4.0 number traces to one results page.

  • Results page: benchmark/results/2026-09-add-4.0-vs-vanilla.md consolidates pilot rounds 3–5 (mutation, edges, halts, evidence reruns, cost, beat shares, rules that did and didn't transfer).
  • Book chapters: 03, 05, 06, 08, 19, 20, the glossary, the worked example and the checklists state the rules the benchmark asked for:
    • checks send inputs the way a real caller does;
    • silent authority reads least-privilege;
    • every RULES id gets a check;
    • the report puts the costliest assumption first;
    • subagents only for security work.
  • Measured sections: both READMEs and ch 20 gain a measured-on-4.0 section; the CHANGELOG 4.0.0 entry gains a Measured section.
  • Animated page: add-method/docs/add-value.html is a self-contained page that honours reduced motion and has a table view and a noscript summary.

ADD task

docs-4-value: freeze, verify PASS.

Test plan

  • python3 -m pytest -q add-method/tests/test_docs_value.py: 9 checks
  • mkdocs build --strict

Added: ADD vs vanilla Claude, as one request run both ways (task docs-value-tradeoff)

  • Both READMEs, rewritten value half:
    • Highlights;
    • what ADD gives your project;
    • the same request run both ways (a side-by-side flow table from two real round-6 runs);
    • measured on Sonnet 5.5, with every gain priced at 1.7–2.1× the dollars and 4.0–5.1× the minutes, and vanilla's ties and lead named;
    • when vanilla is the right call.
  • New page add-method/docs/add-vs-vanilla.html:
    • plays both transcripts on one clock: vanilla done at 46 s for $0.28; ADD sealed at 97 s, verified at 186 s, done at 195 s for $0.63;
    • then shows where the decisions went, and a scoreboard of the n = 3 means;
    • self-contained, with reduced-motion and no-JS states, a table view, and #t= to open at a moment.
  • Round 6 carried into: the results page, ch 20 and the CHANGELOG.
  • python3 -m pytest -q add-method/tests: 154 passed · mkdocs build --strict ✓

…very 4.0 number to one results page

Contract and red doc checks: no page sends non-security work to a subagent or fans out a lane;
the Direction and report pages state the input-shape, least-privilege, every-rule-covered and
costliest-first rules; a results page consolidates benchmark rounds 3-5 with n and disclosures;
README, ch 20 and the animated value page quote only numbers found on it, cost beside gain.

author: Tin Dang
… shows what ADD buys Claude Code

The book still sent data and architecture work to a fresh subagent (ch 05, ch 20, the glossary,
the checklist, the worked example) and told Explore and parallel work to fan out subagents; it now
states the merged budget: one foreground subagent per beat, for security work only. Ch 03, ch 06
and the checklists gain the rules the skill carries: every RULES id covered, inputs sent the way a
real caller sends them, least privilege for a silent authorization, the costliest guess reported
first.

benchmark/results/2026-09-add-4.0-vs-vanilla.md consolidates benchmark rounds 3-5 with n,
disclosures and links to each pilot report. Both READMEs, ch 20 and the CHANGELOG quote only its
numbers, cost beside gain. add-method/docs/add-value.html is a self-contained animated tour of
those numbers (validated palette, light and dark, reduced motion, a no-JS summary, a table view),
linked from both READMEs and the book home. The strict book build passes.

author: Tin Dang
Seal intact against 2a40874; the doc checks (9), the add-method suite (148, with the front-door
honesty guards) and the benchmark suite (522) are green, and the strict book build publishes the
value page. Probes: every visible number on the page traces to the results page, and each cost
multiple matches its dollars. The page was rendered and read in light, dark, narrow, animated and
reduced-motion modes; four layout faults and a blank no-JS page were fixed on sight.

author: Tin Dang
…s by beat is refuted

On claude-sonnet-5-5 the same ADD 4.0 skill ran Direction in 1.1-1.7 min against 4.1 on Sonnet 5,
and cost 1.7-2.1x vanilla against 2.8-2.9x. A Haiku main session with a Sonnet advisor never
consulted it in 6 runs, cost 1.6-1.8x add-4, and failed every wm1 app on a non-stdlib framework;
the lean skill never handed Build to a Haiku subagent. The lean edits that cut work (fewer stubs,
persona by grep) trimmed wall time 14-38% with no quality loss. On this model both workloads
saturate on edges and mutation; ADD keeps its lead on the ambiguous spec. n = 3, not significant.

author: Tin Dang
Comment thread add-method/tests/test_docs_value.py Fixed
…aude as one request run both ways

Both READMEs will open with what ADD gives a project, then run one request both ways (vanilla and
ADD) as a side-by-side mockup, animated on a new book page from two real round-6 runs. Round 6 on
Sonnet 5.5 moved the numbers: vanilla now ties on held-out edges and mutation, so every gain is
priced in dollars and minutes and the ties are named. Six checks fail red; the docs-4-value checks
stay green as the regression.

author: Tin Dang
…age plays both flows

Both READMEs now pitch ADD by what it leaves in a project: a contract per change, every guess on
the record, least-privilege readings of silence, checks sealed before the code, re-runnable
evidence, and memory that outlives the chat. The same request then runs both ways, from two real
round-6 runs on Sonnet 5.5: vanilla done at 46 s for $0.28 with its choices in the chat reply; ADD
done at 195 s for $0.63 with the same choice recorded as an assumption and cancel sealed owner-only.

The measured section prices every gain at 1.7-2.1x the dollars and 4.0-5.1x the minutes, and names
where vanilla ties (edges, mutation) and where it leads (surfacing the contradiction). A new page,
docs/add-vs-vanilla.html, plays both transcripts on one clock (reduced-motion and no-JS states, a
table view, #t= to open at a moment). The results page, ch 20 and the CHANGELOG carry round 6.

author: Tin Dang
The build commit landed at 138 s and the probe on a free port at 154 s belongs to Verify, as the
animated page already shows.

author: Tin Dang
Seal intact; 6 checks pass fresh, the package suite (154) and the loop census (27) are green, the
strict docs build passes, and the flow page was rendered and read at desktop, mid-run, dark and
narrow widths.

author: Tin Dang
@TinDang97 TinDang97 changed the title docs: the book matches the merged 4.0 skill, and an animated page shows what ADD buys Claude Code docs: the book matches 4.0, and the READMEs trade ADD against vanilla Claude with an animated side-by-side flow Sep 30, 2026
Comment thread add-method/tests/test_docs_tradeoff.py Fixed
… tags in any case

CodeQL (py/bad-tag-filter) flagged the regex that removes <script>/<style> blocks before the
brag and source-run checks read the visible text: it matched lower case only. It now strips any-case
tags with attributes and a spaced closing tag, in both this task's check and docs-4-value's C8.
Stricter, not weaker; intent unchanged. 154 package tests pass.

author: Tin Dang
@TinDang97
TinDang97 merged commit 90cdc8a into main Sep 30, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants