Your AI's first milestone is always great. ADD is for every milestone after that.
Describe the feature. The agent plans, seals, builds and verifies it — and hands you a report of every decision it made, backed by evidence you can re-run.
- 📜 Every change leaves its reasoning in your repo — the rules, the guesses and the checks live in one task file next to the code, so the next session or teammate reads the intent instead of guessing it.
- 🛡️ Safer guesses where your spec is silent — ADD takes the least-privilege reading when nobody said who may act, and seals a check for it.
- 🔒 Trust rests on evidence, not a plausible diff — checks are sealed in a
freezecommit before the build, and every task ends in a verdict you can re-run. Security findings always lead the report. - ⚖️ An honest price — Measured on ADD 4.0.0 with Claude Sonnet 5.5, n = 3 per arm per workload: 1.7–2.1× the dollars and 4.0–5.1× the minutes of vanilla Claude Code. The code quality ties; ADD gets more of a spec's silences right (results).
- 💸 Ceremony only where it buys trust — most changes take the Quick lane: one red→green test, one commit, no task file.
Vanilla Claude Code already writes good code: on Sonnet 5.5 it passed every held-out edge case in our benchmark. What it does not leave behind is the why. ADD 4.0 is one skill file, with no engine and no CLI, that makes every change leave the reasoning in your repository, next to the code:
| What your project keeps | Why it pays off later | |
|---|---|---|
| 📜 A contract per change | .add/tasks/<slug>.md: the rules, the guesses, and the checks that prove them |
the next session, or the next teammate, reads the intent instead of reverse-engineering it from code |
| 🙋 Every guess on the record | each silence in your request becomes an ASSUMPTION: the reading taken and the cost if wrong, costliest first | you review a short list of decisions, not a diff; a wrong guess is one line to correct |
| 🛡️ Safer readings of silence | when nobody said who may do something, ADD takes the least-privilege reading and seals a check for it | the booking benchmark never said who may cancel: vanilla let anyone in 8 of 8 runs; ADD chose owner-only in 3 of 3 on Sonnet 5.5 |
| 🔒 Checks sealed before the code | the failing checks are committed as freeze(<slug>) before any build |
a test weakened to get green shows up in git diff; a change of intent is a visible refreeze commit |
| 🔬 Evidence you can re-run | a verdict (PASS, RISK-ACCEPTED or HARD-STOP) with the exact commands and counts, committed as verify(<slug>) |
the claimed test count matched a fresh rerun in 33 of 33 benchmark runs; security findings always lead the report |
| 🧠 Memory that outlives the chat | state lives on disk, not in the conversation | over six evolving milestones, one long chat's requirement coverage fell .92 → .75, while fresh sessions resuming from disk held 1.0 (report) |
Two real round-6 runs on Sonnet 5.5, given the same spec: a booking service whose requirements
contradict each other (a conflict gets 202 waitlisted and 409 rejected) and say nothing on six
other decisions. ▶ Watch both play on one clock.
| clock | 🏃 vanilla Claude Code (vanilla-amb/rep1) |
🛡️ Claude Code + ADD (add-4-amb/rep1) |
|---|---|---|
| 0 s | reads the repo | orients from .add/PROJECT.md |
| 31 s | writes app/ and 7 tests in one command |
… |
| 46 s | tests pass, smoke run OK: done, $0.28 | … |
| 56–94 s | Direction: writes .add/tasks/booking-waitlist.md with 17 rules, 8 ASSUMPTIONS and 19 checks, plus the tests |
|
| 97 s | runs the 19 checks red and seals them: freeze(booking-waitlist) |
|
| 110–138 s | Build: code until all 19 pass, sealed files untouched, then a build commit | |
| 154–186 s | Verify: a live probe with curl, a fresh run (19 OK) and an empty seal diff, then verify(booking-waitlist): PASS |
|
| 195 s | done, $0.63 | |
| the contradiction | caught it and chose waitlist-by-default with a "waitlist": false opt-out, in its chat reply |
caught it and made the same choice, as ASSUMPTION A1 in the task file, first in the report as the costliest guess |
| who may cancel a booking? | anyone | the owner only (rule R:OWNER, sealed check C11) |
| left in your repo | the code and its tests | the code, the checks, the contract and the evidence, as three commits |
Same code quality, same headline decision. The difference is where that decision lives, and the guesses nobody asked about.
The same Claude Code with and without the ADD skill, on claude-sonnet-5-5 (Sonnet 5.5), n = 3 per
arm per workload (results ·
both flows, animated ·
the earlier rounds, animated).
| on Sonnet 5.5 (wm1 · amb1) | vanilla Claude Code | + ADD 4.0 |
|---|---|---|
| you pay: dollars per run | $0.31 · $0.33 | $0.52 · $0.68, 1.7× · 2.1× |
| you pay: minutes per run | 0.8 · 0.85 min | 4.1 · 3.4 min, 5.1× · 4.0× |
| held-out edge cases passed | 19 of 19 · 14 of 14 | 19 of 19 · 14 of 14, a tie |
| seeded bugs its own tests catch (mutation) | 0.83 · 0.76 | 0.75 · 0.79, within noise |
| planted ambiguities handled right, of 7 (amb1) | 4.3 | 5.7 |
| "who may cancel?" read as owner-only (amb1) | 0 of 3 | 3 of 3 |
| surfaced the spec's contradiction (amb1) | 2 of 3 | 1 of 3 |
| claimed test count = a fresh rerun | no claim made | 6 of 6 |
On the older Sonnet 5 (rounds 4–5), ADD's own tests caught more seeded bugs (0.68 vs 0.51 and 0.79 vs 0.53) and vanilla shipped no tests in 2 of 5 runs, at 2.2–2.9× the dollars. Sonnet 5.5 closed those gaps on these workloads, and ADD's cost fell from $1.67 to $0.52 a run.
Fine print: "vanilla" is Claude Code carrying the operator's own ~/.claude config
(which already asks for red/green TDD), not bare Claude Code. Both workloads are saturated at
Sonnet 5.5, and n = 3 is direction, not proof.
- Throwaway work (a script, a spike, a one-shot): use vanilla. It runs 4–5× faster at about half the price, and on a strong model the code is as good.
- A product you will still be changing next month (several milestones, teammates, or anything where who may do this? matters): use ADD. The decisions outlive the chat, the guesses get reviewed, and the checks cannot quietly weaken.
- In between, ADD sizes each request itself. Most changes take the Quick lane, which is one red→green test and one commit, with no task file.
ADD works with the agent you already use (Claude Code, Codex, Cursor, Copilot, Gemini) and installs via npm, pip, or the Claude Code plugin.
Direction before speed. Trust comes from evidence you can re-run, not from reading code and finding it plausible.
Prerequisites: Node ≥ 18 (npm path) or Python ≥ 3.10 (pip path), a git repository, and a coding agent.
npx @pilotspace/add init # Node / npmpip install pilotspace-add && pilotspace-add init # Python / pip# Claude Code plugin — no npm or pip needed
/plugin marketplace add pilotspace/ADD
/plugin install add@add-method
See a real one: this repo's own
.add/folder.
In Claude Code, run /add and say what you want to build:
/add 'Let users log in with email + password / SSO, and keep them signed in for 30 days unless they explicitly log out.'The agent:
- 🧭 Orients from
.add/PROJECT.mdand the open task files — never re-reading your whole repo. - 📐 Sizes the request: Quick, Task, Explore or Milestone.
- ✍️ Writes the direction — rules, assumptions and failing checks — and seals them in a
freezecommit. - ✅ Builds and verifies to a written verdict, then reports; a security finding always goes to the top.
/add status | continueState lives on disk, not in the chat.
One task · three beats · one file. Every change worth a contract is one task file at
.add/tasks/<slug>.md. Direction writes its rules, assumptions and checks, runs the
checks red, and seals them with a freeze(<slug>) commit. Build turns them green without
touching the sealed files. Verify proves the seal is intact, runs everything fresh, reads
what tests cannot show, tries to break the green, and writes one verdict — PASS,
RISK-ACCEPTED, or HARD-STOP — committed as verify(<slug>). The decisions are what you keep;
the code is disposable.
Tasks compound into milestones; milestones grow the project. A milestone lists its tasks up front, runs each just in time, and is done only when every exit criterion carries its evidence.
- 📖 Read the book — the full method, chapter by chapter
- 🆕 What changed in 4.0 — the engine removed, and why
- ⚖️ ADD vs spec-kit — the honest comparison
- ⚡ Getting Started · 🔍 Full walkthrough
- 📒 Beyond code — a month-end close, end to end
- 📊 Benchmark results — every trust and cost claim, reproducible from this repo
- 📦 Package source · Changelog
- 🗞️ ADD Across the Org: AI-Driven Development Beyond Code
Releases: @pilotspace/add (npm) · pilotspace-add (PyPI)
MIT License · pilotspace/ADD

