ADD 4.0 — one skill, no engine - #226
Merged
Merged
Conversation
SKILL.md (174 lines) is now the whole method: orient from .add/ and git, size the work into a lane (Quick, Task, Explore, Milestone), then drive Direction (rules, assumptions, failing checks) sealed by a freeze(<slug>) commit, Build to green inside the seal, and Verify by diffing the sealed files against the freeze, running the checks fresh, reading the residue, refuting, and writing one verdict into the task's EVIDENCE. There is no human approval gate; the session report lists every assumption taken so the human can review after. references/format.md defines the hand-maintained ABF-1 subset (PROJECT, specs, milestones, tasks, personas); references/explore.md the research lane. The 3.x phase guides and helper scripts are removed. The three shipped skill trees stay byte-identical; tests/test_skill_only.py guards the shape (one skill, no engine, no CLI instruction, git as the seal). BREAKING CHANGE: the skill no longer calls the `add` CLI; every 3.x verb is retired. author: Tin Dang
Delete add-method/tooling/ (add.py, cli.py, templates, pins), the add-worker/add-advisor roster in every tree, the bundled engine copy, FORMAT.md, and the engine-only scripts. Remove tests/engine, tests/skill (3.x prose pins) and the covers-grammar test: what they held no longer exists. The nine starter personas move out of tooling/templates to add-method/personas/ as plain persona files (type: Persona, title:, sources:) that orient on the task file and git instead of CLI verbs. Version parity drops the ENGINE declaration (eight remain). CI no longer materialises a vendored engine before the suite. BREAKING CHANGE: `add` CLI, `.add/tooling/`, and the agent roster are gone. author: Tin Dang
Move this repo's 3.x bundle (145 tasks, 35 milestones, specs, personas) to archive/add-3x-bundle/ beside the 2.x record, and seed a 4.0 bundle: PROJECT.md with four invariants and the test command, five fresh specs whose decisions cite the 2026-09-28 ratification and the benchmark record, the add-4-skill-only milestone, and three personas (method-steward and feature-builder carried and restated for 4.0, plus the security-reviewer starter). engine-notary and gate-security-reviewer guarded engine surfaces that no longer exist and stay archived. The dogfood test now checks the 4.0 shape; .gitignore drops the engine cache entries; SECURITY.md names 4.0.x as the supported line. author: Tin Dang
…ing else Both installer twins (bin/cli.js 348 lines, _installer.py 347 + _cli.py) replace ~4k lines and drop the @clack/prompts dependency. A project install refreshes the skill and the vendored persona corpus by stage-then-swap, seeds the starter personas and PROJECT.md without ever overwriting, writes the managed ADD block (new begin marker, legacy one recognised) into CLAUDE.md and AGENTS.md, refreshes a stale 3.x block in .clinerules and similar files only where one exists, and last removes a 3.x .add/tooling/ and ADD's own roster agents when they are recognisably ours. --global installs the skill for the user. --help prints help; unknown flags exit 2 without writing. Packaging ships personas/ and no engine; prepare_bundle.py regenerates _bundled/ for 4.0. publish.yml and teacher-refresh.yml no longer call the engine; CI no longer installs npm deps; marketplace and dependabot text updated. Installer tests rewritten red-green: 89 pass, including a real upgrade over published 3.6.0 artifacts. BREAKING CHANGE: interactive prompts and the removed flags (--force, --no-skill, --stage, prune-data) are gone. author: Tin Dang
Rewrite the book and front-door docs for the skill-only method. Every chapter keeps its idea and swaps the mechanism: the freeze is a freeze(<slug>) commit, the gate is a verdict written into EVIDENCE, the receipt is real command output, and the human reviews after through a report that lists every assumption. The command reference becomes 13 · Files and commits; the bundle chapter condenses format.md; the personas chapters merge; a new chapter 20 maps every 3.x mechanism to its 4.0 equivalent, gives upgrade steps, and says plainly what is given up (mechanical refusal). Appendix D is a real run whose first refute caught a vacuous concurrency check and shows the refreeze. CLAUDE.md, AGENTS.md and .clinerules carry one 4.0 block under the new installer marker. The 3.x book tests and the engine-driven fixture are replaced by tests/book/test_book.py (nav, links, no retired verbs outside the migration page, format parity, worked-example loop) and test_beyond_code.py; shipped-docs, claim-truth and orientation tests are updated. Agent-support claims now match what the installer writes (AGENTS.md and CLAUDE.md). mkdocs build --strict passes. author: Tin Dang
Bump every version declaration to 4.0.0 (package.json, package-lock.json, pyproject.toml, __init__.py, plugin.json; the three SKILL.md trees already read 4.0.0). CHANGELOG [4.0.0] records what changed, what was removed, the installer's new surface, and the 3.x migration path. The three manifest descriptions now describe the skill-only method. Close the add-4-skill-only milestone: all five EXIT criteria ticked with their evidence. Full suite: 128 passed; mkdocs build --strict passes. Publishing stays with the human: merge after the 3.7.0 PR, then tag. author: Tin Dang
A single-file page (React + Framer Motion from esm.sh, no build step) showing what changed between 3.7 and 4.0: the layer stack with and without the engine, one Task's steps and engine calls, the loop stage by stage, the values kept / mechanisms changed / machinery removed, measured size deltas as paired small-multiple bars with a table view, the 3.x benchmark cost, and the trade-off 4.0 makes (mechanical refusal for git-checkable claims). Every figure is measured from the two commits or cited from benchmark/. The two series colors pass the palette validator in light and dark; reduced motion is honored. author: Tin Dang
Add two arms for a head-to-head pilot. add-4 installs this branch's add-method with the 4.0 installer; add-3x installs the release/3.7.0 worktree, refusing to run unless its HEAD is the pinned fba5445. Both give each workspace its own git repo and baseline commit, so 4.0's freeze/verify commits land in the workspace, not the harness repo, and 3.7's gate has the working tree it requires. The add-skill prompt wrapper is add-loop clause for clause, with the method mechanics swapped (read SKILL.md; freeze/verify commits; verdict in EVIDENCE), and a test holds its length within 15%. add-loop gains the `brief` step 3.7's gate requires. loop_census.py records, from real workspace git state, task files, freeze/refreeze/verify commits, EVIDENCE verdicts, seal integrity and red-before-seal, and reports measured: false rather than a vacuous zero when the workspace cannot be read. The broken `add` arm is retired with a clear refusal; its archived records still score. A test guard fails any test that would reach the real claude CLI. benchmark/tests: 506 passed, 12 skipped. author: Tin Dang
….7 pilot loop_census.red_first missed unittest failures wrapped in ANSI color codes: the codes broke `FAILED (failures=` and the word boundary before `AssertionError`, so a real red run before the seal was not counted. Strip ANSI codes before matching; a test replays the pilot's real output. PILOT-4v3-2026-09-28.md records the first head-to-head (n=1 per cell, same model): quality and loop adherence held for 4.0 on wm1 and amb1, cost was level to 21% higher, total tokens level to 42% lower. Not significant; the next campaign is 3 reps over wm1-3 and amb1. author: Tin Dang
The 4.0 pilot put the method's cost in turns, not bytes: every turn re-reads the whole context, and 4.0 ran about 7 turns over a no-method run on wm1 (21 vs 14; 680k vs 420k tokens). The extra turns were a read of references/format.md for the task shape, separate red-run and seal turns, separate evidence writing, and fix-ups. SKILL.md now carries the task template inline, so a Task needs no second file; orients in one command; and adds a Turns section: Direction is two turns (write task + tests together; one command runs the checks and commits the freeze), Build batches files, Verify is two turns (one command for seal diff + checks + regression; one for EVIDENCE + the verify commit). No step is dropped — only round-trips. 197 lines. Guard tests hold the inline template and the turn rule. author: Tin Dang
…llows the batched loop The turn-budget rerun (2 reps × wm1, amb1) cut 4.0 from 21 to ~15.5 turns and 680k to ~490k tokens on wm1 (no-method floor: 14 / 420k) with quality unchanged and every seal intact. One run in four sealed after a red caused only by an import error; the run that avoided it wrote NotImplementedError stubs first. SKILL.md now says so in the Direction turn rule (198 lines). The census missed three shapes the batched loop produces, each fixed with a test from the real output: a test run inside the seal command now counts toward red-first; the seal is the task file plus the files its CHECKS name, not everything riding in the freeze commit; and a stub's NotImplementedError counts as a right-reason red. PILOT-4v3-2026-09-28.md records the floor, the anatomy and round 2. author: Tin Dang
…nvariant and the persona-routing roadmap author: Tin Dang
… rules The 4.0 cut dropped the enforcement behind several closed-loop invariants (ADD_3_6_Closed_Loop_Research.md §17) and adopted none of the persona roadmap (ADD_Dynamic_Persona_Research.md). Each gap is now a line on the path the model walks, with no engine and no extra turns on ordinary work: - rules name their source or are marked derived:; checks name a falsifier - a task names its risks:, which pick the persona, a second evidence mode and the residue lenses - floor work: a second reader under the counter-lens before the seal (replacing the removed human pre-approval) and executable counter-lens probes after the build - Verify runs the consumers of a changed gives: surface - Quick work that touches the floor is a Task now - tag only verified work; observes: names what to watch after release - an escape opens a successor (fixes:) that closes only on a bound prevention; zero-yield controls become method deltas Depth lives in two new references read on a trigger: evidence.md (the loop from intent to production) and personas.md (routing, counter-lens, lens: traces, evals, lifecycle). Every starter and repo persona gains covers-risks, evidence and counter-lens. SKILL.md stays at 200 lines. author: Tin Dang
…issue leads the report Skill-creator eval, iteration 1 (3 fixtures, new skill vs the pre-change snapshot, claude-sonnet-5) found two regressions in the first cut: - quick-escalation: asked for a typo fix on a log line that writes the user's API token, the new skill filed the leak as an "open risk" under "HARD-STOPs: None"; the old skill led with it as a HARD-STOP. The tripwire only covered the change itself. Now a security issue met in passing stays out of the diff but leads the report as a HARD-STOP. - gives-consumer: a consumed-surface task (floor) skipped the second reader and wrote no lens: line; the invite task skipped the subagent "for this size". The second reader is now every floor task however small: a fresh subagent for security, data and architecture, a cold reread under the counter-lens otherwise; lens: is always written, or "none — why". Iteration 2: every assertion held on all three fixtures (29/29 vs 23/29 for the old skill). The invite run's pre-seal second reader added a rule neither earlier run had: an invite is not consumed if add_member fails. author: Tin Dang
Seal intact (freeze 837097b, 0-line diff to head 4e591e8); on the clean tree the guard file passes 13/13 and the full suite 134/134. Six mutation probes on a temp copy were each caught by their guard. The skill-creator eval (3 fixtures, new skill vs the pre-change snapshot) found two regressions in iteration 1, fixed in 4e591e8; iteration 2 held every assertion (29/29 vs 23/29). Extra cost appears only on security-floor work (+46% tokens on invite-expiry), recorded as a method delta to measure. author: Tin Dang
… investigations skill-creator's description loop (claude-opus-5-5, 12 train / 8 held-out queries, 3 runs each; 10 should-trigger, 10 near-miss negatives such as "git add everything", "add a last_login column", writing a lone test file, a PRD, a CI lint fix): - previous description: train 11/12, test 8/8, precision 100%. Its one miss was "investigate why the nightly export slowed down ... cited findings before anyone changes code" — the compressed description had lost the Explore lane. - adopted description: train 12/12, test 8/8, precision 100%. It names evidence-first investigations and says when not to trigger. SKILL.md stays at 200 lines: three sentences that repeated a rule stated elsewhere were folded (the record-real-output line, the floor-only note in Turns, the consumers wrap). author: Tin Dang
…I clone CI checks out shallow with no tags, so the PREVIOUS_VERSION selection found no older tag and test_previous_tag_is_strictly_older failed on every run (py3.10 and py3.12, PR #224 and #226) while passing locally, where the clone has every tag. The test now runs the selection in a throwaway repo holding v3.3.0, v3.4.0, v3.5.0, v3.6.0 and v3.10.0; the last one fails a lexical sort, so the check still proves numeric ordering. author: Tin Dang
…cle cannot see author: Tin Dang
…unts 8 annotation slots The fixture's clamp has two mutants no test can kill, so a strong suite tops out at 6/8; C4 now demands a 0.5 gap between strong and weak instead of an unreachable 0.8. C6's expected ratio was an arithmetic slip (3/8, not 3/6). author: Tin Dang
…ation, static, security, tests, evidence benchmark/quality.py scores a run's workspace, never writing to it: held-out edge suites for wm1 (19 cases) and amb1 (14), validated against a correct and a sloppy reference app; a seeded AST mutation score of the run's own tests (n/a on a red baseline); AST static quality; security smells; test quality; and, for ADD runs, whether the EVIDENCE regression count matches a fresh rerun. `python -m benchmark.quality <runs…>` writes quality.json per run. author: Tin Dang
… real runs write Scoring the 09-29 runs crashed on a `.venv/bin/python` claim (the scoring copy omits .venv) and read n/a on wrapped and unittest-style claims. C10 pins those shapes. author: Tin Dang
…t and reruns the whole suite Never replays the agent's own command: it named a .venv the scoring copy omits and crashed the scorer. author: Tin Dang
Seal intact since the last refreeze; check 12/12, benchmark suite 522 passed. Scoring a real run twice is deterministic and leaves the workspace byte-identical. Round-3 results recorded in benchmark/PILOT-4v3-2026-09-29.md. author: Tin Dang
…y-only subagents, input robustness author: Tin Dang
…gents only for security From benchmark round 3 and the task-contract audit: two contracts left a Must with no check; found: was almost never used; the second reader cost 2.3x on amb1 (three subagents in one run, a wakeup wait in another); five of nine apps crashed on a null or number body. SKILL.md now asks that every RULES id sit on a covers: line, that cheap guesses be checked now, that input surfaces get a malformed-input check, and caps subagents at one per beat, in the foreground, for security work only. Still 200 lines. author: Tin Dang
…ence, explore.md included A refute probe found references/explore.md still telling the agent to "give parallel subagents disjoint questions", against SKILL.md's one-per-beat, foreground budget. The sealed rule R:CONSISTENT named only evidence.md and personas.md, so its check could not see the conflict. R:CONSISTENT now covers every reference and C2 walks REFERENCES; the check is red until explore.md is brought in line. author: Tin Dang
references/explore.md still told the agent to give parallel subagents disjoint questions, against SKILL.md's one-per-beat, foreground budget. Explore now splits disjoint questions and answers them in turn, with at most one foreground subagent for a question whose reading would flood the context. Mirrored to the bundled and project skill trees. author: Tin Dang
…only costs Same-day add-4 vs vanilla, wm1 + amb1, n = 3 each, scored on every quality dimension, with each ADD practice mapped to what it moved. The falsifier-per-check practice is the measured value (mutation score +0.17 / +0.26); ASSUMPTIONS change no decision; the malformed-input rule did not transfer (field-level only); Direction is 46% of tokens and about forty turns against the skill's "two". Discloses the 403 reruns, an estimated lost attempt cost, the operator config both arms load, and one escaped timezone defect. author: Tin Dang
Seal intact against the refreeze 84a0563; the check (15) and the full add-method suite (136) are green on the clean tree, and the three skill trees are identical. Round 4 shows every rule covered in 5 of 6 contracts, found: in every task, and no subagents in 6 of 6 runs. The malformed-input rule was written into every contract but only at field level, so body-level garbage still crashed 4 of 6 apps — accepted and handed to review as a proposed "shape" sweep. A refute probe found explore.md contradicting the subagent budget; refrozen and fixed. author: Tin Dang
The round-4 cost ledger counted a message's thinking, text and each tool call as separate turns, reporting Direction at 37–47 turns and 46% of tokens. Deduplicated by message id it is about 20 messages and 41–44%, with Verify at 2–5%. A traced run shows where Direction's messages go: environment probing one command at a time, batched writes, and a fumbled seal commit. The method delta carries the corrected figure. author: Tin Dang
… callers' inputs, guesses that fail safe Contract and red guard tests for the round-4 findings: Direction as a batched plan (ground and find the test command in one command, parallel writes in one message, red and seal in one); checks that send the body itself malformed and every allowed value form; a least-privilege reading for silent authorization; ASSUMPTIONS reported costliest-if-wrong first. Round 5, same day, is the behavioural evidence. author: Tin Dang
…nputs, guesses that fail safe Direction is now a batched plan: one command grounds and finds how the tests run (installing what is missing), one message writes the task file, tests and stubs as parallel writes, and one command runs red and seals. Checks send inputs the way a real caller does: every value form the spec allows (a timestamp with and without an offset) and a malformed body as well as malformed fields. A silence about who may act or see takes the least-privilege reading, and the report lists the assumptions costliest-if-wrong first. The red-first sentence joined the Seal and the residue lenses point at evidence.md, keeping SKILL.md at 200 lines. Mirrored to all three trees. author: Tin Dang
Seal intact against b4dbf6d; the check (18) and the full add-method suite (139) are green and the three skill trees are identical. Round 5, same day, shows the input-shape rule changed behaviour: no garbage-body crash in 6 of 6 runs (was 4 of 6), offset-aware timestamps in most tests, every wm1 edge case held. The least-privilege reading, the costliest-first report and the three-turn Direction were stated but mostly not followed: prose that advises moves the model less than an artifact it must write. Vanilla wrote no tests in 2 of 6 runs and halted on the contradictory spec in 1 of 3; ADD delivered and tested in 6 of 6. author: Tin Dang
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ADD 4.0 — one skill, no engine
Stacked on #224 (3.7.0). Merge #224 first, then retarget this PR to
main.ADD becomes one markdown skill the model follows with git and the project's own test command. The Python engine, the
addCLI, and theadd-worker/add-advisorroster are removed. Decisions ratified in the 2026-09-28 interview:Why
benchmark/BENCHMARK.md), 3.x ADD cost 18.2M tokens (~$15) against spec-kit's 3.8M ($2.53), at fidelity minimums of 0.97 vs 0.95.benchmark/FINDINGS-2026-08-10.md).What changed (6 commits)
98ccf487SKILL.md(174 lines) plusreferences/format.mdandreferences/explore.md: orient → size (Quick/Task/Explore/Milestone) → Direction →freeze(<slug>)commit → Build → Verify (seal diff · fresh green · residue · refute · verdict in## EVIDENCE) → learn → report0c94030ftooling/, the roster, the engine and 3.x prose tests,FORMAT.md(−62,828 lines). Starter personas move toadd-method/personas/01ec1e2aarchive/add-3x-bundle/; a fresh 4.0.add/gets PROJECT, five specs, the milestone and three personas3014c021c157c3f76a097175What a 3.x upgrade does:
.add/tooling/and ADD's own agents;Evidence
cd add-method && python3 -m pytest -q→ 128 passedmkdocs build --strict→ exit 0--global;Review these first
CLAUDE.md.bak/AGENTS.md.bakwhen it rewrites a managed block (same as 3.x).CLAUDE.mdandAGENTS.md, and adds Gemini's settings entry only if.gemini/exists. It no longer writes.clinerules; I assumed Cline readsAGENTS.mdand did not verify it.benchmark/andblog/. The root imagesadd-install.png,add-task-growth-wheel.pngandadd-milestone-task-lifecycle.pngare no longer referenced;add-install.pngshows the old "approve once" message.After merge (you)
Tag
v4.0.0frommain; publish runs on the tag.author: Tin Dang