Skip to content

ADD 4.0 — one skill, no engine - #226

Merged
TinDang97 merged 35 commits into
mainfrom
feat/add-4-skill-only
Sep 30, 2026
Merged

TinDang97 merged 35 commits into
mainfrom
feat/add-4-skill-only

Conversation

@pilotspacex-byte

Copy link
Copy Markdown
Contributor

ADD 4.0 — one skill, no engine

Stacked on #224 (3.7.0). Merge #224 first, then retarget this PR to main.

ADD becomes one markdown skill the model follows with git and the project's own test command. The Python engine, the add CLI, and the add-worker/add-advisor roster are removed. Decisions ratified in the 2026-09-28 interview:

  • no engine
  • git commit as the seal
  • no approval gate; the human reviews after
  • the ABF-1 subset maintained by hand
  • personas and the teacher corpus kept

Why

  • On the repo's own benchmark (benchmark/BENCHMARK.md), 3.x ADD cost 18.2M tokens (~$15) against spec-kit's 3.8M ($2.53), at fidelity minimums of 0.97 vs 0.95.
  • On amb1, vanilla prompting cost $0.61 against ADD's $2.13–2.52 (benchmark/FINDINGS-2026-08-10.md).
  • Most of the engine's recent growth guarded its own stamps.

What changed (6 commits)

commit what
98ccf487 Skill. SKILL.md (174 lines) plus references/format.md and references/explore.md: orient → size (Quick/Task/Explore/Milestone) → Direction → freeze(<slug>) commit → Build → Verify (seal diff · fresh green · residue · refute · verdict in ## EVIDENCE) → learn → report
0c94030f Engine removed. tooling/, the roster, the engine and 3.x prose tests, FORMAT.md (−62,828 lines). Starter personas move to add-method/personas/
01ec1e2a Dogfood. The 3.x bundle is archived to archive/add-3x-bundle/; a fresh 4.0 .add/ gets PROJECT, five specs, the milestone and three personas
3014c021 Installers. Both twins drop to ~350 lines with no dependencies; they upgrade a 3.x project (details below)
c157c3f7 Book. Rewritten for 4.0. New chapter 20 maps every 3.x mechanism to its 4.0 equivalent; Appendix D is a real run
6a097175 Release prep. Versions at 4.0.0, CHANGELOG, milestone closed

What a 3.x upgrade does:

  • refreshes the skill, and ADD blocks in legacy pointer files where they exist;
  • removes .add/tooling/ and ADD's own agents;
  • never overwrites user files.

Evidence

  • cd add-method && python3 -m pytest -q → 128 passed
  • mkdocs build --strict → exit 0
  • Installer smoke:
    • npm and pip give identical trees on a fresh install, a re-run and --global;
    • a real upgrade over the published 3.6.0 wheel/npm tarball passes, with state preserved and the engine removed.

Review these first

  • What 4.0 gives up: mechanical refusal. A dishonest agent could now write a PASS. What replaces the refusal is that every claim is checkable with git: the seal is a diff, and EVIDENCE names exact commands at an exact commit. Chapter 20 says this plainly.
  • No approval gate, and nothing interrupts. A security finding becomes a HARD-STOP verdict that leads the session report; it no longer pauses for a human.
  • 3.x upgrade leaves backups. It writes CLAUDE.md.bak / AGENTS.md.bak when it rewrites a managed block (same as 3.x).
  • Agent support. The installer writes CLAUDE.md and AGENTS.md, and adds Gemini's settings entry only if .gemini/ exists. It no longer writes .clinerules; I assumed Cline reads AGENTS.md and did not verify it.
  • Six 3.x direction tasks were never built. Their red tests were marked xfail in Release 3.7.0 — merge loop-that-closes as-is, the last engine release #224 and are deleted with the engine here.
  • Left as-is: the historical benchmark/ and blog/. The root images add-install.png, add-task-growth-wheel.png and add-milestone-task-lifecycle.png are no longer referenced; add-install.png shows the old "approve once" message.

After merge (you)

Tag v4.0.0 from main; publish runs on the tag.

author: Tin Dang

SKILL.md (174 lines) is now the whole method: orient from .add/ and git,
size the work into a lane (Quick, Task, Explore, Milestone), then drive
Direction (rules, assumptions, failing checks) sealed by a freeze(<slug>)
commit, Build to green inside the seal, and Verify by diffing the sealed
files against the freeze, running the checks fresh, reading the residue,
refuting, and writing one verdict into the task's EVIDENCE. There is no
human approval gate; the session report lists every assumption taken so
the human can review after.

references/format.md defines the hand-maintained ABF-1 subset (PROJECT,
specs, milestones, tasks, personas); references/explore.md the research
lane. The 3.x phase guides and helper scripts are removed. The three
shipped skill trees stay byte-identical; tests/test_skill_only.py guards
the shape (one skill, no engine, no CLI instruction, git as the seal).

BREAKING CHANGE: the skill no longer calls the `add` CLI; every 3.x verb
is retired.

author: Tin Dang
Delete add-method/tooling/ (add.py, cli.py, templates, pins), the
add-worker/add-advisor roster in every tree, the bundled engine copy,
FORMAT.md, and the engine-only scripts. Remove tests/engine, tests/skill
(3.x prose pins) and the covers-grammar test: what they held no longer
exists.

The nine starter personas move out of tooling/templates to
add-method/personas/ as plain persona files (type: Persona, title:,
sources:) that orient on the task file and git instead of CLI verbs.
Version parity drops the ENGINE declaration (eight remain). CI no longer
materialises a vendored engine before the suite.

BREAKING CHANGE: `add` CLI, `.add/tooling/`, and the agent roster are gone.

author: Tin Dang
Move this repo's 3.x bundle (145 tasks, 35 milestones, specs, personas)
to archive/add-3x-bundle/ beside the 2.x record, and seed a 4.0 bundle:
PROJECT.md with four invariants and the test command, five fresh specs
whose decisions cite the 2026-09-28 ratification and the benchmark
record, the add-4-skill-only milestone, and three personas
(method-steward and feature-builder carried and restated for 4.0, plus
the security-reviewer starter). engine-notary and gate-security-reviewer
guarded engine surfaces that no longer exist and stay archived.

The dogfood test now checks the 4.0 shape; .gitignore drops the engine
cache entries; SECURITY.md names 4.0.x as the supported line.

author: Tin Dang
…ing else

Both installer twins (bin/cli.js 348 lines, _installer.py 347 + _cli.py)
replace ~4k lines and drop the @clack/prompts dependency. A project
install refreshes the skill and the vendored persona corpus by
stage-then-swap, seeds the starter personas and PROJECT.md without ever
overwriting, writes the managed ADD block (new begin marker, legacy one
recognised) into CLAUDE.md and AGENTS.md, refreshes a stale 3.x block
in .clinerules and similar files only where one exists, and last removes
a 3.x .add/tooling/ and ADD's own roster agents when they are
recognisably ours. --global installs the skill for the user. --help
prints help; unknown flags exit 2 without writing.

Packaging ships personas/ and no engine; prepare_bundle.py regenerates
_bundled/ for 4.0. publish.yml and teacher-refresh.yml no longer call
the engine; CI no longer installs npm deps; marketplace and dependabot
text updated. Installer tests rewritten red-green: 89 pass, including a
real upgrade over published 3.6.0 artifacts.

BREAKING CHANGE: interactive prompts and the removed flags (--force,
--no-skill, --stage, prune-data) are gone.

author: Tin Dang
Rewrite the book and front-door docs for the skill-only method. Every
chapter keeps its idea and swaps the mechanism: the freeze is a
freeze(<slug>) commit, the gate is a verdict written into EVIDENCE, the
receipt is real command output, and the human reviews after through a
report that lists every assumption. The command reference becomes
13 · Files and commits; the bundle chapter condenses format.md; the
personas chapters merge; a new chapter 20 maps every 3.x mechanism to
its 4.0 equivalent, gives upgrade steps, and says plainly what is given
up (mechanical refusal). Appendix D is a real run whose first refute
caught a vacuous concurrency check and shows the refreeze.

CLAUDE.md, AGENTS.md and .clinerules carry one 4.0 block under the new
installer marker. The 3.x book tests and the engine-driven fixture are
replaced by tests/book/test_book.py (nav, links, no retired verbs
outside the migration page, format parity, worked-example loop) and
test_beyond_code.py; shipped-docs, claim-truth and orientation tests are
updated. Agent-support claims now match what the installer writes
(AGENTS.md and CLAUDE.md). mkdocs build --strict passes.

author: Tin Dang
Bump every version declaration to 4.0.0 (package.json, package-lock.json,
pyproject.toml, __init__.py, plugin.json; the three SKILL.md trees
already read 4.0.0). CHANGELOG [4.0.0] records what changed, what was
removed, the installer's new surface, and the 3.x migration path. The
three manifest descriptions now describe the skill-only method.

Close the add-4-skill-only milestone: all five EXIT criteria ticked with
their evidence. Full suite: 128 passed; mkdocs build --strict passes.
Publishing stays with the human: merge after the 3.7.0 PR, then tag.

author: Tin Dang
A single-file page (React + Framer Motion from esm.sh, no build step)
showing what changed between 3.7 and 4.0: the layer stack with and
without the engine, one Task's steps and engine calls, the loop stage by
stage, the values kept / mechanisms changed / machinery removed, measured
size deltas as paired small-multiple bars with a table view, the 3.x
benchmark cost, and the trade-off 4.0 makes (mechanical refusal for
git-checkable claims). Every figure is measured from the two commits or
cited from benchmark/. The two series colors pass the palette validator
in light and dark; reduced motion is honored.

author: Tin Dang
Add two arms for a head-to-head pilot. add-4 installs this branch's
add-method with the 4.0 installer; add-3x installs the release/3.7.0
worktree, refusing to run unless its HEAD is the pinned fba5445. Both
give each workspace its own git repo and baseline commit, so 4.0's
freeze/verify commits land in the workspace, not the harness repo, and
3.7's gate has the working tree it requires.

The add-skill prompt wrapper is add-loop clause for clause, with the
method mechanics swapped (read SKILL.md; freeze/verify commits; verdict
in EVIDENCE), and a test holds its length within 15%. add-loop gains
the `brief` step 3.7's gate requires. loop_census.py records, from real
workspace git state, task files, freeze/refreeze/verify commits,
EVIDENCE verdicts, seal integrity and red-before-seal, and reports
measured: false rather than a vacuous zero when the workspace cannot be
read. The broken `add` arm is retired with a clear refusal; its archived
records still score. A test guard fails any test that would reach the
real claude CLI.

benchmark/tests: 506 passed, 12 skipped.

author: Tin Dang
….7 pilot

loop_census.red_first missed unittest failures wrapped in ANSI color
codes: the codes broke `FAILED (failures=` and the word boundary before
`AssertionError`, so a real red run before the seal was not counted.
Strip ANSI codes before matching; a test replays the pilot's real output.

PILOT-4v3-2026-09-28.md records the first head-to-head (n=1 per cell,
same model): quality and loop adherence held for 4.0 on wm1 and amb1,
cost was level to 21% higher, total tokens level to 42% lower. Not
significant; the next campaign is 3 reps over wm1-3 and amb1.

author: Tin Dang
The 4.0 pilot put the method's cost in turns, not bytes: every turn
re-reads the whole context, and 4.0 ran about 7 turns over a no-method
run on wm1 (21 vs 14; 680k vs 420k tokens). The extra turns were a read
of references/format.md for the task shape, separate red-run and seal
turns, separate evidence writing, and fix-ups.

SKILL.md now carries the task template inline, so a Task needs no second
file; orients in one command; and adds a Turns section: Direction is two
turns (write task + tests together; one command runs the checks and
commits the freeze), Build batches files, Verify is two turns (one
command for seal diff + checks + regression; one for EVIDENCE + the
verify commit). No step is dropped — only round-trips. 197 lines.
Guard tests hold the inline template and the turn rule.

author: Tin Dang
…llows the batched loop

The turn-budget rerun (2 reps × wm1, amb1) cut 4.0 from 21 to ~15.5
turns and 680k to ~490k tokens on wm1 (no-method floor: 14 / 420k) with
quality unchanged and every seal intact. One run in four sealed after a
red caused only by an import error; the run that avoided it wrote
NotImplementedError stubs first. SKILL.md now says so in the Direction
turn rule (198 lines).

The census missed three shapes the batched loop produces, each fixed
with a test from the real output: a test run inside the seal command
now counts toward red-first; the seal is the task file plus the files
its CHECKS name, not everything riding in the freeze commit; and a
stub's NotImplementedError counts as a right-reason red.
PILOT-4v3-2026-09-28.md records the floor, the anatomy and round 2.

author: Tin Dang
…nvariant and the persona-routing roadmap

author: Tin Dang
… rules

The 4.0 cut dropped the enforcement behind several closed-loop invariants
(ADD_3_6_Closed_Loop_Research.md §17) and adopted none of the persona
roadmap (ADD_Dynamic_Persona_Research.md). Each gap is now a line on the
path the model walks, with no engine and no extra turns on ordinary work:

- rules name their source or are marked derived:; checks name a falsifier
- a task names its risks:, which pick the persona, a second evidence mode
  and the residue lenses
- floor work: a second reader under the counter-lens before the seal
  (replacing the removed human pre-approval) and executable counter-lens
  probes after the build
- Verify runs the consumers of a changed gives: surface
- Quick work that touches the floor is a Task now
- tag only verified work; observes: names what to watch after release
- an escape opens a successor (fixes:) that closes only on a bound
  prevention; zero-yield controls become method deltas

Depth lives in two new references read on a trigger: evidence.md (the
loop from intent to production) and personas.md (routing, counter-lens,
lens: traces, evals, lifecycle). Every starter and repo persona gains
covers-risks, evidence and counter-lens. SKILL.md stays at 200 lines.

author: Tin Dang
…issue leads the report

Skill-creator eval, iteration 1 (3 fixtures, new skill vs the pre-change
snapshot, claude-sonnet-5) found two regressions in the first cut:

- quick-escalation: asked for a typo fix on a log line that writes the
  user's API token, the new skill filed the leak as an "open risk" under
  "HARD-STOPs: None"; the old skill led with it as a HARD-STOP. The
  tripwire only covered the change itself. Now a security issue met in
  passing stays out of the diff but leads the report as a HARD-STOP.
- gives-consumer: a consumed-surface task (floor) skipped the second
  reader and wrote no lens: line; the invite task skipped the subagent
  "for this size". The second reader is now every floor task however
  small: a fresh subagent for security, data and architecture, a cold
  reread under the counter-lens otherwise; lens: is always written, or
  "none — why".

Iteration 2: every assertion held on all three fixtures (29/29 vs 23/29
for the old skill). The invite run's pre-seal second reader added a rule
neither earlier run had: an invite is not consumed if add_member fails.

author: Tin Dang
Seal intact (freeze 837097b, 0-line diff to head 4e591e8); on the clean
tree the guard file passes 13/13 and the full suite 134/134. Six mutation
probes on a temp copy were each caught by their guard. The skill-creator
eval (3 fixtures, new skill vs the pre-change snapshot) found two
regressions in iteration 1, fixed in 4e591e8; iteration 2 held every
assertion (29/29 vs 23/29). Extra cost appears only on security-floor work
(+46% tokens on invite-expiry), recorded as a method delta to measure.

author: Tin Dang
… investigations

skill-creator's description loop (claude-opus-5-5, 12 train / 8 held-out
queries, 3 runs each; 10 should-trigger, 10 near-miss negatives such as
"git add everything", "add a last_login column", writing a lone test file,
a PRD, a CI lint fix):

- previous description: train 11/12, test 8/8, precision 100%. Its one miss
  was "investigate why the nightly export slowed down ... cited findings
  before anyone changes code" — the compressed description had lost the
  Explore lane.
- adopted description: train 12/12, test 8/8, precision 100%. It names
  evidence-first investigations and says when not to trigger.

SKILL.md stays at 200 lines: three sentences that repeated a rule stated
elsewhere were folded (the record-real-output line, the floor-only note in
Turns, the consumers wrap).

author: Tin Dang
…I clone

CI checks out shallow with no tags, so the PREVIOUS_VERSION selection found
no older tag and test_previous_tag_is_strictly_older failed on every run
(py3.10 and py3.12, PR #224 and #226) while passing locally, where the
clone has every tag. The test now runs the selection in a throwaway repo
holding v3.3.0, v3.4.0, v3.5.0, v3.6.0 and v3.10.0; the last one fails a
lexical sort, so the check still proves numeric ordering.

author: Tin Dang
Brings d0af5bb from feat/loop-that-closes so the stack stays linear for
review. test_persona_instruction_truth.py stays deleted (4.0 removed the
engine suites); test_candidate_artifact_smoke.py keeps the 4.0 copy, which
already carries the same scratch-repo tag fix (aae677c).

author: Tin Dang
…unts 8 annotation slots

The fixture's clamp has two mutants no test can kill, so a strong suite tops out at 6/8; C4 now demands a 0.5 gap between strong and weak instead of an unreachable 0.8. C6's expected ratio was an arithmetic slip (3/8, not 3/6).

author: Tin Dang
…ation, static, security, tests, evidence

benchmark/quality.py scores a run's workspace, never writing to it: held-out edge suites for wm1 (19 cases) and amb1 (14), validated against a correct and a sloppy reference app; a seeded AST mutation score of the run's own tests (n/a on a red baseline); AST static quality; security smells; test quality; and, for ADD runs, whether the EVIDENCE regression count matches a fresh rerun. `python -m benchmark.quality <runs…>` writes quality.json per run.

author: Tin Dang
… real runs write

Scoring the 09-29 runs crashed on a `.venv/bin/python` claim (the scoring copy omits .venv) and read n/a on wrapped and unittest-style claims. C10 pins those shapes.

author: Tin Dang
…t and reruns the whole suite

Never replays the agent's own command: it named a .venv the scoring copy omits and crashed the scorer.

author: Tin Dang
The CI fix on PR #224 (d0af5bb, test-only) moved the release-3-7 worktree off fba5445, so the pin check refused the add-3x arm. The engine code is unchanged; the pin follows the commit that will ship as 3.7.0.

author: Tin Dang
Seal intact since the last refreeze; check 12/12, benchmark suite 522 passed. Scoring a real run twice is deterministic and leaves the workspace byte-identical. Round-3 results recorded in benchmark/PILOT-4v3-2026-09-29.md.

author: Tin Dang
…y-only subagents, input robustness

author: Tin Dang
…gents only for security

From benchmark round 3 and the task-contract audit: two contracts left a Must with no check; found: was almost never used; the second reader cost 2.3x on amb1 (three subagents in one run, a wakeup wait in another); five of nine apps crashed on a null or number body. SKILL.md now asks that every RULES id sit on a covers: line, that cheap guesses be checked now, that input surfaces get a malformed-input check, and caps subagents at one per beat, in the foreground, for security work only. Still 200 lines.

author: Tin Dang
…ence, explore.md included

A refute probe found references/explore.md still telling the agent to "give parallel subagents
disjoint questions", against SKILL.md's one-per-beat, foreground budget. The sealed rule
R:CONSISTENT named only evidence.md and personas.md, so its check could not see the conflict.
R:CONSISTENT now covers every reference and C2 walks REFERENCES; the check is red until
explore.md is brought in line.

author: Tin Dang
references/explore.md still told the agent to give parallel subagents disjoint questions,
against SKILL.md's one-per-beat, foreground budget. Explore now splits disjoint questions and
answers them in turn, with at most one foreground subagent for a question whose reading would
flood the context. Mirrored to the bundled and project skill trees.

author: Tin Dang
…only costs

Same-day add-4 vs vanilla, wm1 + amb1, n = 3 each, scored on every quality dimension, with each
ADD practice mapped to what it moved. The falsifier-per-check practice is the measured value
(mutation score +0.17 / +0.26); ASSUMPTIONS change no decision; the malformed-input rule did not
transfer (field-level only); Direction is 46% of tokens and about forty turns against the skill's
"two". Discloses the 403 reruns, an estimated lost attempt cost, the operator config both arms
load, and one escaped timezone defect.

author: Tin Dang
Seal intact against the refreeze 84a0563; the check (15) and the full add-method suite (136)
are green on the clean tree, and the three skill trees are identical. Round 4 shows every rule
covered in 5 of 6 contracts, found: in every task, and no subagents in 6 of 6 runs. The
malformed-input rule was written into every contract but only at field level, so body-level
garbage still crashed 4 of 6 apps — accepted and handed to review as a proposed "shape" sweep.
A refute probe found explore.md contradicting the subagent budget; refrozen and fixed.

author: Tin Dang
The round-4 cost ledger counted a message's thinking, text and each tool call as separate turns,
reporting Direction at 37–47 turns and 46% of tokens. Deduplicated by message id it is about 20
messages and 41–44%, with Verify at 2–5%. A traced run shows where Direction's messages go:
environment probing one command at a time, batched writes, and a fumbled seal commit. The
method delta carries the corrected figure.

author: Tin Dang
… callers' inputs, guesses that fail safe

Contract and red guard tests for the round-4 findings: Direction as a batched plan (ground and
find the test command in one command, parallel writes in one message, red and seal in one);
checks that send the body itself malformed and every allowed value form; a least-privilege
reading for silent authorization; ASSUMPTIONS reported costliest-if-wrong first. Round 5, same
day, is the behavioural evidence.

author: Tin Dang
…nputs, guesses that fail safe

Direction is now a batched plan: one command grounds and finds how the tests run (installing what
is missing), one message writes the task file, tests and stubs as parallel writes, and one command
runs red and seals. Checks send inputs the way a real caller does: every value form the spec
allows (a timestamp with and without an offset) and a malformed body as well as malformed fields.
A silence about who may act or see takes the least-privilege reading, and the report lists the
assumptions costliest-if-wrong first. The red-first sentence joined the Seal and the residue
lenses point at evidence.md, keeping SKILL.md at 200 lines. Mirrored to all three trees.

author: Tin Dang
Seal intact against b4dbf6d; the check (18) and the full add-method suite (139) are green and the
three skill trees are identical. Round 5, same day, shows the input-shape rule changed behaviour:
no garbage-body crash in 6 of 6 runs (was 4 of 6), offset-aware timestamps in most tests, every
wm1 edge case held. The least-privilege reading, the costliest-first report and the three-turn
Direction were stated but mostly not followed: prose that advises moves the model less than an
artifact it must write. Vanilla wrote no tests in 2 of 6 runs and halted on the contradictory spec
in 1 of 3; ADD delivered and tested in 6 of 6.

author: Tin Dang
@TinDang97
TinDang97 changed the base branch from feat/loop-that-closes to main September 30, 2026 06:08
@TinDang97
TinDang97 marked this pull request as ready for review September 30, 2026 06:08
@TinDang97 TinDang97 closed this Sep 30, 2026
@TinDang97 TinDang97 reopened this Sep 30, 2026
@TinDang97
TinDang97 merged commit ae2c87d into main Sep 30, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants