docs: the book matches 4.0, and the READMEs trade ADD against vanilla Claude with an animated side-by-side flow - #231
Merged
Merged
Conversation
…very 4.0 number to one results page Contract and red doc checks: no page sends non-security work to a subagent or fans out a lane; the Direction and report pages state the input-shape, least-privilege, every-rule-covered and costliest-first rules; a results page consolidates benchmark rounds 3-5 with n and disclosures; README, ch 20 and the animated value page quote only numbers found on it, cost beside gain. author: Tin Dang
… shows what ADD buys Claude Code The book still sent data and architecture work to a fresh subagent (ch 05, ch 20, the glossary, the checklist, the worked example) and told Explore and parallel work to fan out subagents; it now states the merged budget: one foreground subagent per beat, for security work only. Ch 03, ch 06 and the checklists gain the rules the skill carries: every RULES id covered, inputs sent the way a real caller sends them, least privilege for a silent authorization, the costliest guess reported first. benchmark/results/2026-09-add-4.0-vs-vanilla.md consolidates benchmark rounds 3-5 with n, disclosures and links to each pilot report. Both READMEs, ch 20 and the CHANGELOG quote only its numbers, cost beside gain. add-method/docs/add-value.html is a self-contained animated tour of those numbers (validated palette, light and dark, reduced motion, a no-JS summary, a table view), linked from both READMEs and the book home. The strict book build passes. author: Tin Dang
Seal intact against 2a40874; the doc checks (9), the add-method suite (148, with the front-door honesty guards) and the benchmark suite (522) are green, and the strict book build publishes the value page. Probes: every visible number on the page traces to the results page, and each cost multiple matches its dollars. The page was rendered and read in light, dark, narrow, animated and reduced-motion modes; four layout faults and a blank no-JS page were fixed on sight. author: Tin Dang
…s by beat is refuted On claude-sonnet-5-5 the same ADD 4.0 skill ran Direction in 1.1-1.7 min against 4.1 on Sonnet 5, and cost 1.7-2.1x vanilla against 2.8-2.9x. A Haiku main session with a Sonnet advisor never consulted it in 6 runs, cost 1.6-1.8x add-4, and failed every wm1 app on a non-stdlib framework; the lean skill never handed Build to a Haiku subagent. The lean edits that cut work (fewer stubs, persona by grep) trimmed wall time 14-38% with no quality loss. On this model both workloads saturate on edges and mutation; ADD keeps its lead on the ambiguous spec. n = 3, not significant. author: Tin Dang
…aude as one request run both ways Both READMEs will open with what ADD gives a project, then run one request both ways (vanilla and ADD) as a side-by-side mockup, animated on a new book page from two real round-6 runs. Round 6 on Sonnet 5.5 moved the numbers: vanilla now ties on held-out edges and mutation, so every gain is priced in dollars and minutes and the ties are named. Six checks fail red; the docs-4-value checks stay green as the regression. author: Tin Dang
…age plays both flows Both READMEs now pitch ADD by what it leaves in a project: a contract per change, every guess on the record, least-privilege readings of silence, checks sealed before the code, re-runnable evidence, and memory that outlives the chat. The same request then runs both ways, from two real round-6 runs on Sonnet 5.5: vanilla done at 46 s for $0.28 with its choices in the chat reply; ADD done at 195 s for $0.63 with the same choice recorded as an assumption and cancel sealed owner-only. The measured section prices every gain at 1.7-2.1x the dollars and 4.0-5.1x the minutes, and names where vanilla ties (edges, mutation) and where it leads (surfacing the contradiction). A new page, docs/add-vs-vanilla.html, plays both transcripts on one clock (reduced-motion and no-JS states, a table view, #t= to open at a moment). The results page, ch 20 and the CHANGELOG carry round 6. author: Tin Dang
The build commit landed at 138 s and the probe on a free port at 154 s belongs to Verify, as the animated page already shows. author: Tin Dang
Seal intact; 6 checks pass fresh, the package suite (154) and the loop census (27) are green, the strict docs build passes, and the flow page was rendered and read at desktop, mid-run, dark and narrow widths. author: Tin Dang
… tags in any case CodeQL (py/bad-tag-filter) flagged the regex that removes <script>/<style> blocks before the brag and source-run checks read the visible text: it matched lower case only. It now strips any-case tags with attributes and a spaced closing tag, in both this task's check and docs-4-value's C8. Stricter, not weaker; intent unchanged. 154 package tests pass. author: Tin Dang
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The book and READMEs now match the merged ADD 4.0 skill. Every 4.0 number traces to one results page.
benchmark/results/2026-09-add-4.0-vs-vanilla.mdconsolidates pilot rounds 3–5 (mutation, edges, halts, evidence reruns, cost, beat shares, rules that did and didn't transfer).add-method/docs/add-value.htmlis a self-contained page that honours reduced motion and has a table view and a noscript summary.ADD task
docs-4-value: freeze, verify PASS.Test plan
python3 -m pytest -q add-method/tests/test_docs_value.py: 9 checksmkdocs build --strictAdded: ADD vs vanilla Claude, as one request run both ways (task
docs-value-tradeoff)add-method/docs/add-vs-vanilla.html:#t=to open at a moment.python3 -m pytest -q add-method/tests: 154 passed ·mkdocs build --strict✓