Release 3.7.0 — merge loop-that-closes as-is, the last engine release - #224
Merged
Merged
Conversation
The 3.6 engine was read rung by rung against the closed-loop research (design page, 2026-09-11). Seven seams are open and three of them are already claimed in shipped prose: "dependents that need: it are flagged stale" (four files, no reader, a §3.5 FORMAT never wrote), "the regression floor" (one direction.md bullet, no slot, no reader), and a refute rung a builder's own T1 read satisfies at a human floor. This milestone closes them with one new verb (add release), one new FORMAT section and zero new node types; every new refusal arms at the refute rung's arming or higher. Nodes: receipt-anchored-to-head · regression-floor · consumers-go-stale · release-stamp · escape-with-prevention · successor-not-reopen · refute-tier-floor · must-carries-source · quick-lane-tripwire · observes-slot · explores holdout-that-holds and method-health. Milestone frozen at plan authority under the ratified design. author: Tin Dang
A receipt recorded the blobs it observed and not the commit, so nothing in the bundle could say whether a tag shipped the tree a PASS verified. `run` now records `head:` — HEAD when the run STARTED, before the command can move it — and `committed:` — true exactly when every scope_digest blob is the blob HEAD's tree holds at that path, decided by one `ls-tree` over the digest and never by whole-tree cleanliness (R:COMMITTEDBYCLAIM). Outside git, or on an unborn branch, neither key is written and the note names the cause (R:INVENTEDHEAD); a reader treats absence as unknown, so a pre-3.7 receipt is neither releasable nor unreleasable by default. FORMAT §8.1 states the anchor. Four twins mirrored, ENGINE_MD5 re-aimed. Red first: 7 checks failed on the absent keys; full suite 1556 green. Task: .add/tasks/receipt-anchored-to-head.md — gate PASS on receipt 1. author: Tin Dang
…and the gate reads direction.md listed "the regression floor" among what PLAN carries; nothing gave it a grammar, a slot or a reader, and the 3.2 cut shipped a task green over a red host because the floor lived in memory. Now PLAN carries one line — `regression: full | affected · <cmd> · <why>` or `none · <why>` — and two rungs read it, armed exactly where the refute rung arms: `freeze` refuses a rung-bound task with no floor (R:NOFLOOR), and `gate PASS` refuses a declared full|affected floor that was never run, ran stale or ran red (R:FLOORUNRUN), naming which, with the PLAN's own command as the fix. `add run --floor` records the host suite as an ordinary receipt carrying `floor: regression`; `latest_receipt` never returns it (R:FLOORASGATE) and `latest_floor_receipt` answers for it. The verify hint replays the command. FORMAT §8.5 states the grammar; direction.md, verify.md, docs 03 and 05 carry it; the skill surface stays line-neutral (funded by compressing). Fixtures that freeze rung-bound tasks now declare `none · fixture`. Red first: 9 checks failed; full suite green. A fresh-session T2 refute then FOUND two readers that still took a floor receipt for the narrow one — `_latest_run_cid` (the hint demanded a refute the gate never asked for) and the `test_cmd` memory (a floor run became the next build hint); a second read found `_beat_of` (a floor-first run closed the build beat). All three are frozen as E7, E8 and E9 with bound checks, fixed, refrozen and re-read by a third fresh session. This task declared the milestone's first floor line and recorded the first floor receipt (the full suite) before the rung that reads it existed. Task: .add/tasks/regression-floor.md author: Tin Dang
…le, and the consumer's gate holds
Four shipped sentences promised that dependents citing a moved `#gives`
are "flagged stale" and cited a FORMAT §3.5 that was never written; the
refreeze branch wrote one stamp and told nobody. Every freeze stamp now
carries `gives: <digest>` (the published surface alone) and, on a
consumer, `needs: "<target>#gives=<sha8>"` — what it read at its OWN
freeze. Three readers compare pins to the provider's current digest,
digests never dates (R:CLOCKPIN): `doctor` emits one `needs_stale` warn
per (consumer, provider); `todo` hints the verb on the consumer's row;
a rung-bound consumer's `gate PASS` refuses R:STALENEEDS until it
re-crosses. The refreeze that moved a `gives:` names its consumers in its
note. The provider is never blocked by what its consumers pinned
(R:PROVIDERBLOCKED); a stamp with no pin (pre-3.7) answers nothing.
FORMAT §3.5 is written; §8.5 names which hint reader replays the floor;
intake.md, build.md, appendix-c and appendix-d now name the finding they
promised, bound by a check.
Red first: 8 checks failed; full suite green. Four fresh-session T2 reads
each FOUND one defect. Three were one class — a pin property decided on
written text instead of the resolved target: provider order (A5, E6), a
provider named twice (M3, E7), one provider under two spellings (M1, E8);
the writer now dedupes by `_norm` target and readers speak only of OPEN
consumers (E9). The fourth was the unit itself: every reader compared the
pin against the provider's LIVE `gives:` list, so a consumer frozen before
its provider pinned a draft, a silent edit with no refreeze flagged stale,
and a v1->v2->v1 round trip named a consumer whose pin was current. The
unit is now the provider's STAMPED digest (`stamped_gives`): an unfrozen
provider pins `?`, a silent edit moves nothing, and the refreeze note is
the same comparison doctor makes (E10-E12). The fifth found the pin
string itself: `,`/`=`-delimited with no escaping, so a pasted digest in
a `needs:` ref read back as attested. A ref carrying a delimiter now pins
`?` and the reader parses from the right (E13). The sixth found the dedupe
key stripped the fragment, so `#findings` written before `#gives` dropped
the contract pin; the unit is now (resolved node, fragment), and a scalar
`gives:` digests as one surface (E14, E15). The seventh found the reader
still split on `,` before parsing `=`, so a pasted two-entry stamp value
misread as a digest; an unattestable ref is now written with its
delimiters stripped and the reader accepts only exact `<ref>=<sha8|?>`
tokens, proven by a property check over every combination (E16). The
eighth found the pin bypassed the discipline every interpolated value
takes through `_oneline`: a `"` or `{` in a sibling ref corrupted the flow
map and dropped the honest pin or swallowed the next stamp. The stripped
alphabet is now the serializer's own (`_PIN_UNSAFE`), swept by the
property (E17). The ninth found the refreeze note walked LIVE `needs:`
edges while doctor, todo and the gate read the stamp, so an unsealed
`needs:` edit split them; every reader now sources the consumer set from
the stamp's pin (E18). Each fix is frozen with a check; a tenth session
re-read the result.
Filed, not fixed here: the frontmatter BLOCK-list parser swallows the
rest of the frontmatter when an item carries `{` or `'` (the eighth read's
side note) — an authoring hole outside this pin's contract.
Task: .add/tasks/consumers-go-stale.md
author: Tin Dang
…d stops promising a finding E10 pins in silence The tenth T2 read of consumers-go-stale held, with two side notes taken here as a direct change. The gate's fix text forecast `add freeze c, rebuild, add run c` — taken literally it lands on R:UNBRIEFED, because a refreeze re-seals the direction and the gate demands a brief since the last freeze; the recipe now names `add brief <slug>` after the re-cross. Appendix-d promised that a `needs:` pointing at a never-frozen `gives:` "surfaces as an `edge_unresolved` finding" — it does not: E10 pins `?` and says nothing, and `edge_unresolved` fires only when the node file is absent. The sentence now says so. One check in test_consumers_go_stale.py binds both (red first); full suite 1590 passed. Twins mirrored, ENGINE_MD5 re-aimed. Lesson filed via `add learn add` with the block-list parser hole the eighth read exposed. author: Tin Dang
…ified it
A PASS proves a source state and, since receipt-anchored-to-head, a
receipt names the commit it observed — but nothing in the bundle could
say which tree a tag shipped, so production telemetry was evidence about
an unknown build. `add release <tag> --milestone m --by who` is the 29th
verb and the one new record of 3.7: it resolves the tag's tree with
read-only git (`rev-parse`, `ls-tree` — never `tag`, `push`, `publish`:
R:OUTWARD), proves that tree holds every scope blob the members' gated
narrow receipts recorded (one mismatch or absence refuses R:UNANCHORED
naming the task, the path and both blobs; a digest-less receipt refuses
by name; a receiptless explore is skipped by name), and appends
`{ act: release, tag, tree, receipts }` to each done milestone. A
milestone not done refuses R:NOTDONE; a tag git cannot resolve,
R:NOSUCHTAG. `--artifact` and `--build` are recorded verbatim and never
verified (R:PROVENANCEJUDGED): provenance is the pipeline's. `status
--all` names the tag at the released row's end; `show <m>` carries one
`act: release` line per stamp. No new node type — the milestone is the
release's subject.
`_tree_blobs` is the one ls-tree reader, now shared with
`_committed_to_head`. FORMAT §8.6 states the stamp, the anchor and the
read-only rule; docs/16 §16.5 and docs/13 carry the verb; loop.md names
it in one clause (skill surface line-neutral). Nine verb registries
re-aimed 28 -> 29. Red first: 8 checks failed; full suite green. A
fresh-session T2 read FOUND the anchor reading the LATEST narrow receipt
— a red run after the PASS moved it onto a FAIL receipt; the anchor is
now the receipt the newest gate stamp cites (E7). It also found a 78-char
tag pushing the status row past ROW_WIDTH; the title yields first, then
the tag is cut (E8). A second read re-entered the same class through
`gate HARD-STOP` after done: the anchor followed the newest gate, and a
HARD-STOP is a finding, not a verdict — the anchor is now the newest
CLOSING gate (PASS or RISK-ACCEPTED), and a milestone in which no member
anchors refuses by name instead of stamping "0 receipts" (E9, E10). A
third read found the success note dropping the not-done members M3
promised, and asked that the closing gate postdate the member's last
reopen — the window `done` reads (E11, E12). Each frozen with a check; a
fourth read held.
Task: .add/tasks/release-stamp.md
author: Tin Dang
…rned The fourth T2 read of release-stamp held, with one note taken here as a direct change: `_anchor` read whatever receipt cid the member's gate stamp cited, so a hand-edited stamp could borrow another task's digest, carry one cid twice, or point outside the bundle. A cited receipt must now live under the member's own `<slug>.d/runs/<n>.md`; anything else refuses R:UNANCHORED naming the cid and the shape it needed. One check binds it (red first); full suite 1605 passed. Twins mirrored, ENGINE_MD5 re-aimed. author: Tin Dang
…ound prevention The drain existed — learn demands evidence, fold --bind|--reject and R:UNDRAINED hold it — but what it drained was a sentence: a production escape folded with nothing bound to stop the next one. `add learn --escape` now demands `--why-missed` and `--prevention "<check|monitor| method|rule> → <ref>"`, refusing R:UNCAUSED naming the missing or malformed part, and writes them as a tail after the evidence clause so the delta grammar, `deltas`, `search` and the persona loop are untouched. `fold` (folded or --bind) refuses R:UNPREVENTED when an escape's prevention resolves to nothing — a bundle address, a RULES/EDGES id, or a repo file — reading every match first so one dangling prevention leaves every sibling open; `--reject` never reads it. loop.md's observe section classifies every production observation EXPECTED · RULE_VIOLATION · SPEC_SILENCE: only the last two file a delta, and a silence routes to Direction as an assumption or edge, never a Must. Red first: 19 checks failed; full suite green. A fresh-session T2 read FOUND the resolver passing on any existing path (`.`, a directory, an empty file part) and the tail reader ungated by `· escape` — a plain lesson quoting the grammar was refused at fold, and a `·` inside a flag shadowed the bound ref. The ref must now be a file or an authored id, the reader reads only an escape's tail, and the reserved delimiter is refused at learn (E7–E9). A second read found `(evidence:` inside a flag defeating the tail split — a dangling escape hid behind it and folded — and a backticked `<…>` in an authored rule read as a template placeholder (92 live RULES lines carry one); both refused or read correctly now, and a file must lie inside the repo (E10–E12). A third read found a hand-edited malformed prevention (`alert →`, `Check →`, no arrow, no ref, or the clause deleted) reading as NO prevention and folding — the reader is now kind-agnostic and every clause must be well-formed and resolve, and the escape marker cannot ride `--evidence` either (E13, E14). A fourth read found a line break in the lesson or the evidence pushing the tail onto a second physical line the rung never reads, so a dangling escape folded at exit 0 through the documented CLI; `learn` now writes every interpolated value on one line, as it already did for the why-missed and the ref (E15). A fifth found the escape rung gated on the VALUE's truthiness rather than the flag's presence — `--why-missed ""` with no `--escape` recorded a plain lesson at exit 0, which is exactly the unmarked escape this task closes; it now reads presence, and a whitespace-only evidence is refused instead of writing a delta the engine's own reader calls `no_evidence` (E16). A sixth found the rung reading one physical line while deltas.md's frozen grammar says continuation lines join into one delta — a hand-wrapped tail folded unprevented, and a wrap before the clause refused with the false reason that no clause existed; `fold` now reads the delta the grammar defines (E17). A seventh read found that widening the unit had opened a second evidence channel: the tail was keyed on the LAST `(evidence: …)`, so an indented continuation citing one pushed the marker out of view and a dangling escape folded at exit 0 — while a note quoting the grammar in a code span was refused as an escape it never was. The reader now anchors on the `· escape` MARKER itself with backticked spans masked, and `learn` refuses the marker in the lesson as it already did in the evidence: the marker is the engine's own and rides in through no flag. The same read found one delta answered three ways — `deltas` called the grammar's own wrap malformed while `open_delta_count` counted it and `fold` read it; all three now read `joined_deltas`, the grammar's unit (E18, E19). An eighth read found the writer and the reader masking different text: `learn` masked each value alone while the rung masks the whole line, so one stray backtick in the lesson and one in the why-missed paired into a span that swallowed the engine's own marker — the tail was written, `deltas` printed it, and `fold` folded (and `--bind` bound) at exit 0 with nothing to see. `learn` now closes an unpaired backtick in every value it interpolates and refuses one in the prevention ref, so the two views are one view; at fold an odd backtick count means a hand edit the writer never balanced, and there the reading that REFUSES wins (E20). A ninth read found the marker BOUNDING the read instead of gating it: `_prevention_of` scanned from the marker on, so a clause standing before it was skipped — and `--evidence` could put one there, folding a delta whose written tail carried a prevention resolving to nothing, while a resolving clause written before the marker was reported absent. The marker now gates only; every clause anywhere in the delta must resolve, which is what A5 always said, and a marker is a marker however it is punctuated — `· escape:` no longer degrades to prose (E21). A tenth read turned the instrument around and found the WRITERS severing the unit the readers had been taught to join: `join` harvested head lines alone, so a wrapped escape its own stream refused arrived in main stripped of its tail and folded there, and `learn` spliced a new head between a wrapped delta and its continuation, folding the escape and handing its tail to an innocent lesson. Both writers now work in the grammar's unit (E22). Its siblings closed with it: a prevention may not name the escape it is bound to, an id resolves only where a rule is AUTHORED — never one merely spelled under `## LESSONS` or quoted inside a fence — and a clause is read however it is written, so `· Prevention:` and a continuation indented with a non-breaking space are refused by name instead of vanishing (E23). An eleventh read turned on the walker that fix had just added: it read `## ` before the fence, so a heading quoted inside a fenced block turned fencing off and the section on — a fence could forge an authored rule for an escape to name, and the same line hid the genuinely authored rules that followed one. The fence is now read first, a quoted heading opens nothing, and a heading that names nothing authors nothing instead of raising: every exit is a refusal or a record (E24). A twelfth read found the odd-count safety guarding only the marker: a stray backtick AFTER it masked one clause of several, and the delta folded beside a dangling ref. The raw view now decides EVERY clause, not just the marker. Its siblings: a fence is a fence in every spelling a markdown grammar allows, a rule line that says nothing authors nothing, and — the one that defeated the promise end to end with no hand edit — two escapes may no longer name each other. A prevention that names an escape still OPEN binds nothing, whether that escape is another or itself (E25). A thirteenth read found two readers of one address disagreeing: `resolve` strips whitespace on both sides of the `#`, the open-escape guard's own regex admitted none, so `rule → /specs/method.md# M1` passed the ref rung and was invisible to the guard — an escape prevented itself and bound a decision citing itself. Every reader of a prevention address now normalises it the one way `resolve` does. Its siblings: a fence opening a list item is a fence, and a label with nothing after it is a clause the author WROTE, reported as malformed rather than absent (E26). A fourteenth read showed the fence rule was still enumerating list markers rather than deciding what a fence IS: an ordered or nested item opening one left the block unfenced, so a quoted rule read as authored and a genuinely authored rule after it vanished. A fence is now its run of backticks or tildes and the closer that matches it, whatever markers or quote arrows carry the opener (E27). A fifteenth read found the fence rule still matching line shapes rather than tracking a block, and caught one of these very checks passing for the wrong reason: a fence's own indented run closed it early, so a quoted rule read as authored and an escape bound a decision citing it, while a blockquoted fence could never close at all and every genuinely authored rule after it refused. A fence is now a block — its closer shares the container, the character, the length, and sits within three spaces of its opener — an ATX heading counts at up to three spaces, and every fence row now proves the fence CLOSED as well as that it hid what it quoted (E28). A sixteenth read mutation-tested the checks themselves and found four of the five clauses of that closer definition unbound — each could be deleted with the suite still green — and then defeated the promise outright: a heading written inside a blockquote or a list item turned the authoring section back on, so an escape naming an id under `## LESSONS` folded and bound a decision at exit 0 through the documented CLI. A heading now opens or ends a section only at document level, a fence ends with the block that carries it, and every clause of the closer is bound by a check that turns red when it is deleted (E29). That made five of sixteen findings land in one markdown walker, so the root moved: `rules_of` and `edges_of` — the readers the GATE binds coverage with — now read through the same walker the fold rung uses. One fact, one reader: the gate can no longer demand coverage of an example the rung calls unauthored (E30). The full suite is green at 1702. A seventeenth read found the fence's own CONTENT closing it again, this time through a list marker or a tab: the closer test compared only quote depth, and indent was counted in characters. Content is opaque now — a line that opens a list item opens a block and closes nothing — and indent counts columns with a tab as four. The same read found `_section_of` deciding for itself what a heading is, so a node whose RULES heading carried one space of indent owed no check for any Must while the fold rung bound preventions to those same Musts; the walker now re-emits the heading it kept, and the slicer stops deciding. Each of the eleven clauses behind those rules turns the suite red when it is deleted (E31). An eighteenth read defeated the promise through three documented verbs: `join` re-mints a colliding stream id in the delta's HEAD only, so an escape whose prevention named its own address went on naming the id it used to have — which in main is a different lesson. The stream refused it, main folded it and bound a decision on it, and the mutual-prevention cycle came back with it. A re-mint now re-points the address the delta's own tail names, in one pass so a chain cannot re-point a line twice. The same read showed the join check exercising the one ref shape a re-mint cannot move, and three frozen clauses with no check at all: the marker's own spacing, the open-escape status, and the id form. All four are bound now, and four guards that could never fire are gone (E32). A nineteenth read showed that fix was both too narrow and too wide: the re-map was scoped to the spec being written, so an escape in one lens kept naming an id that moved in another and main folded and bound a decision on it, while the same substitution ran over every fresh line, so with two streams a tail whose own id never moved was re-pointed at the other stream's unrelated lesson. The merge now mints every spec first, collects one remap PER STREAM, and re-points only that stream's own lines: an address moves in every lens, and only for the stream whose id actually moved (E33). A twentieth read found the merge to be a THIRD reader of a delta address, and a stricter one: it demanded the `#` and the id be contiguous, so a re-mint skipped every spaced spelling the contract blesses, and the escape its own stream refused folded and bound a decision in main. The merge now reads an address the one way `resolve` does and re-emits it canonical, and the join checks are parametrized over the spellings the address check already pinned — two checks that each covered half a rule and never crossed (E34). A twenty-first read found a bare `## ` matching no branch of the section walker at all, so authoring state was INHERITED: an id jotted under one after `## RULES` resolved, and the escape folded and bound a decision. A heading that names nothing now ends the section it follows, while a deeper `###` opens none. The same read showed the check for that very clause sitting after `## LESSONS`, where authoring was already off — it passed with its subject withheld and stayed green with the clause deleted. It now sits where it bites, and the ref's own one-line normalisation, a Must nothing had ever exercised, is bound beside it (E35). A twenty-second read found the section SLICER deciding for itself what a heading is — a second reader, fence-blind, and the prevention rung consults it before the authoring walk is ever reached: a `## my example` spelled only inside a fenced block under `## LESSONS` opened a section an escape could bind to and fold at exit 0, while the `- E9` in that same fence was correctly refused. One walker now answers "is this line a heading" for both readers. The same read mutation-tested the resolver and found three clauses unbound — the heading form, the frontmatter-key form, and the fence closer's own container test, whose row never reached the closer because the auto-close fired first — each bound now by a check that turns red when it is deleted (E36). A twenty-third read found the fix one form short: `resolve` reads a fragment three ways, and only the heading form had been taught the walker. The delta-id reader stayed fence-blind, so ONE fence gave two answers — the quoted `- E9` refused while the quoted delta id beside it folded the escape and bound a decision at exit 0 — and the same blindness made `deltas` list that example as a real open delta for a human to drain. A node now has one view of the text it LIVES: a fence quotes and never authors, for every form. Its siblings came from the same read's mutations: a `#` title and a `###` sub-heading opened a section for the slicer while the authoring walk read both as content, and the empty `scope:`/`verified:` slots resolved — an address the *check that passes on nothing* shape had reached inside the grammar itself. A four-space indented example is recorded as a spec silence answered: a fence is the only code block a node has, because rules are list items and their wraps indent (A7, E37). A twenty-fourth read turned that fix around and found the WRITERS still on raw lines: `learn` located `## Deltas` and its insertion point by a raw scan, so a spec whose Deltas section opens with a fenced example of the grammar took the new delta INSIDE the fence — filed at exit 0 and thereafter invisible to `deltas`, `search`, `status`, `doctor` and `fold`, which is worse than folding unprevented: nobody is ever asked to drain it. A writer now finds its section the one way `_section` does, and the insert scan walks past a quoted example instead of stopping in it. Two siblings fell with it: `_bind_decision` was a fourth heading reader, so `--bind` reported a bound decision at exit 0 into a `###` or fenced heading the brief could not see; and the merge read main's holdings from raw lines, so a delta main merely QUOTED looked already held and the stream's refused escape was dropped. `--bind` now writes under `learn`'s own two laws — one physical line, and the engine's marker cannot ride in through a flag, which had forged a real open escape (A8). The empty slot rule reaches every spelling, not just the empty list, and a fence raises no mint floor. Every clause is bound by a check that runs the WRITER after planting the fence and asserts the reader sees what was written — eight mutations, eight reds (E38). A twenty-fifth read found the one writer whose CALLER's value leads the line: `learn` writes its own `[COMP · id · status · date]` head first, so no flag can ever open a block, but `--bind` writes the sentence at column zero — and one leading with a fence run opened a real block that blanked the rest of the spec. The dangling escape left `deltas`, `search` and `status`, the engine wrote `open_deltas: 0` itself, `fold` answered R:NOMATCH forever, and the next escape filed landed inside that fence under a second `## Deltas` heading, at exit 0 throughout. A sentence leading with a delta head forged an open delta nobody filed. Naming those shapes would have frozen the fix at the two the read happened to find, so the engine now READS BACK what it is about to write: a decision may not change the delta ids or the sections the spec's own readers see. The same read reproduced the note the read before it raised — `add init` writes no `tasks/`, so a first `join` raised where every exit must be a refusal or a record; it records (E39). A twenty-sixth read turned the read-back on itself: it compared the delta ids and the heading NAMES — what the guard happened to know — and never the decision lines its own readers read. So a decision carrying a bare `<tenant>` landed, was reported bound, and every reader disowned it: `brief` rendered the section `unauthored="true"`, `doctor` called it scaffold, and the NEXT bind deleted it without a word, because the placeholder detector cannot tell the engine's seed line from a decision the engine itself just wrote. The writer now refuses a sentence its readers would disown, and the read-back reaches the decision lines, with the pure seed line the one exemption. Its mirrors: `(evidence:` is the engine's own marker in EVERY value, not the two an earlier read named — `_delta_identity` split a lesson at the FIRST one, so two streams whose lessons differed only past it collapsed into one false conflict and `join` dropped both refused escapes; identity now keys on the last marker as every other reader does. And the explore gate read Musts raw while `rules_of` read them authored, so it demanded a finding for an `M9` quoted inside a fence the fold rung calls unauthored — E30's own law, one reader further out (E40). A twenty-seventh read found the rung that never got a read-back: a fence that NEVER closes — the case every fence edge before it had frozen about what CLOSES one — made `learn` report the id it minted at exit 0 while the line landed inside the fence. Both writer branches lose there: the insert scan walks past a fence that never ends to EOF, and a blanked heading reads as "this spec has no Deltas", so the second branch appends one inside the same fence. `deltas` said none, `search` found nothing, `fold` answered R:NOMATCH forever, `show` answered R:NOSUCHNODE for the id `learn` had just handed back, and the engine wrote `open_deltas: 0` itself. The primary escape writer now reads back what `--bind` already did — the line it wrote must be the line the readers read — and refuses R:UNREADABLE naming the spec. Its siblings: `_authored_rules` re-emitted the author's heading verbatim while `_section_of` matches it exactly, so a `## RULES (frozen)` left the gate seeing no Musts while the fold rung resolved ids under it; the walker re-emits the section it RECOGNISED. And the merge was the one delta writer that never recomputed `open_deltas:` (E41). A twenty-eighth read found the sibling writer that never got that read-back: on the very spec where `learn` now refuses, `fold --bind` reported a bound decision at exit 0 while the EOF branch wrote the decision AND the heading it created inside the same never-closing fence. Its guard asked only that nothing was LOST, and under a fence all three of those comparisons compare empty to empty and pass — so the guard that protected the decision could not see the decision. The bind now asks `learn`'s question too: is what it WROTE addressable. Its siblings, all from the same read: a prevention ref carrying a NUL reached the filesystem and raised where every exit must be a refusal or a record; a byte-order mark before a spec's frontmatter raised an AttributeError where `doctor` already names the file; and M1's "split on the FIRST arrow" was a greedy `\S+` that backtracked to the LAST one whenever no space separated them, refusing a genuine `check->…` with a message naming a kind nobody wrote — the spaced control is why the bound check missed it (E42). A twenty-ninth read named the pattern the three reads before it had been walking: the read-back was bound to the writers each read happened to touch, never to "a writer of a delta line". The third one — the merge, the writer that carries an escape BETWEEN bundles — had no guard at all, so joining into a spec whose `## Deltas` sits under a never-closing fence exited 0 and filed the stream's refused escape where `deltas`, `status`, `doctor` and `fold` could never see it, on the very spec where `learn` and `fold --bind` refuse. The same function raised on a byte-order mark AFTER copying the stream's node — the partial merge R:PHANTOMSTREAM exists to forbid — and `fold` itself, the rung this task is about, raised there for all three of its dispositions. There is now ONE reader of "can a writer land in this spec", asked by the merge before any node is copied, by `fold` for every disposition, and by `learn` (E43). A thirtieth read found the fourth writer — and it is the one every R:UNREADABLE refusal's own `next:` sends the author to. `doctor --sync` recomputed `open_deltas: 1 → 0` from the blinded body at exit 0, erasing the last trace of an escape `fold` had just refused, and on a byte-order mark it raised after rewriting an earlier spec. `doctor` itself named the COUNTER — `delta_count_drift` — and never the blindness beneath it, so the way out the engine points at did not exist. Now doctor names the spec no writer can land in, the drift it cannot honestly count is not reported beside it, and `--sync` repairs nothing there and says what it left. The same read caught E43's `fold` clause frozen, unimplemented and unbound: the guard read `if why and node["raw"] is None`, which makes the fence branch dead, and the check that looked like it covered the clause parametrized the three DISPOSITIONS while exercising one CAUSE. Both halves are bound now, and a join that would carry nothing is no longer refused over a spec it would never touch (E44). A thirty-first read found the pre-flight reading the wrong side: the escape lives in the STREAM's spec, and a never-closing fence there makes the harvester yield nothing — indistinguishable from a stream that filed nothing — so `join` reported `joined 1 stream(s) · specs union-merged` at exit 0 while the escape its own stream had refused was gone, with the worktree about to be discarded. A byte-order mark on the same file carried it through correctly, so the two causes disagreed. Both sides are read now. Two more from the same read: the pre-flight asked "is there a gated node" while the merge asks "…and not HARD-STOP", so one rejected stream — the security case — blocked a whole wave's join over a spec it would never touch; one predicate answers that now. And E42's read-back asked only whether its OWN line was addressable, which a decision that opens a fence at the end of a spec answers yes: it sat readable inside the fence it had just opened, and the next `learn` refused forever. The question is the spec's, not the line's — asked before and after the write (E45). A thirty-second read found the root under three of those fixes: the markdown walkers rebuilt the caller's own `splitlines(keepends=True)` list with a `\n` join, and Python splits on eight more line boundaries than that. One U+2028, form feed or file separator in a spec's prose re-split a line, every index the walkers hand back shifted by one, and `learn` spliced its new head between a wrapped escape's head and its continuation: the dangling escape folded at exit 0 while an innocent lesson refused in its place, and the merge carried the orphan tail alone — `join` reporting success with the stream's refused escape gone. No guard could see it, because the cause is neither a fence nor a missing frontmatter. The walk now reads the list its caller split, and the separator axis is exercised over the BODY the walkers index rather than over the flags `learn` already normalises (E46). The thirty-third read found the marker read four ways. `ESCAPE_MARK` carried no `re.I` while its sibling `PREVENTION_TAIL` does, and three more sites spelled the same regex by hand, so one byte — `· escape` to `· Escape` — made `_prevention_of` return None: `fold --bind` folded the dangling escape AND bound a decision at exit 0, a correctly-spelled escape naming the capitalised one folded too, and `learn` and `--bind` would still let a forged marker ride in through a flag. The marker is now ONE object every site asks, case-blind like the clause below it. The read also severed a tail: a continuation indented with a character Python splits on (form feed, U+2028, VT, FS) leaves the head at column 0 with no marker, so `fold` now refuses a tail that belongs to no delta rather than folding the head beside it — `delta_spans` is the one reader of which lines a delta HOLDS, and `joined_deltas` and `orphan_tail` both read through it (E47, A9). Its two secondary findings are bound too: the join pre-flight's "does this stream file into THIS spec" term, which no check held, and E46's per-line heading walk, which only the E45 check executed. The thirty-fourth read broke that fix. `TAIL_MARK` was anchored at `^`, which enumerated the SHAPE the read that named it produced — a wrap immediately before `· escape` — instead of asking the property: does this stranded line CARRY an escape's tail. Wrapping one word earlier, the shape E17 blesses, left the orphan starting with a word and the head folded at exit 0; a plain unindented newline was enough. Worse, the MERGE had no such guard at all: `_delta_lines` rebuilds only the lines `delta_spans` holds, so `join` dropped the orphan and main folded a laundered, marker-less delta — the escape its own stream had refused. The predicate now asks the property, reads inside `## Deltas` and nowhere else (A10, so prose that mentions a clause blocks no fold), and the merge asks what `fold` asks before it copies anything (E48). Its check sweeps the wrap POSITION, which is the axis the last one held fixed. Twins mirrored, both pins re-aimed, three skill trees line-neutral. Task: .add/tasks/escape-with-prevention.md author: Tin Dang
`verified[]` is append-only, so a reopen keeps the old PASS — but loop.md called the residual case, reopening a task inside an already-closed milestone, "surfaced by `add status --check` as incoherent and resolved by hand". No such finding ever existed: the engine reopened it silently and the milestone's goal-gate, which had counted that task done when it published its exit boxes, was rewritten underneath. `reopen` now refuses a done task whose milestone is `done` or `archived` (R:CLOSEDHISTORY), naming the milestone, its status, and the successor form with the old cid filled in — `add new Task <slug>-2 --supersedes /tasks/<slug>.md`. The rung sits LAST, after the type, `done` and beat checks, so a malformed call is still named by the flag the operator typed, and a milestone under any other word (`shipped`, active, none at all) reopens exactly as before. `new --supersedes <ref>` gives that form a writer: the ref may be a bare slug or a cid, it is RESOLVED before it is written, and a ref that names no node — or names a lesson inside one — refuses R:PHANTOMPREDECESSOR and writes nothing. `supersedes` was one of three edge keys FORMAT counted with zero live uses; `show` already walks every EDGE_KEY, so the successor renders `↓ supersedes` and the predecessor `↑ supersedes` with no renderer change. The old node, its receipts and its PASS are never touched by either path (R:REWRITTENPAST). One collision, resolved by ORDER: release-stamp's frozen M3/E12 needs a member reopened AFTER its closing gate inside a milestone that is done at release time. Its checks now reach that state legally — reopen while the milestone is active, then close it — so the older claim, its assertion and this refusal all stand (E6). Neither side was weakened. loop.md and FORMAT.md say so in place, line-neutral, across all three skill trees. Four twins mirrored, three pins re-aimed. Task: .add/tasks/successor-not-reopen.md author: Tin Dang
…es, and the gate reads the tier
Four beats of milestone `loop-that-closes`, interleaved in one engine file and so in one commit.
must-carries-source — a Must now ends `(from: <where you were told> · fails-on: <the plausible
wrong reading>)`. The interview asks for the source of every Must that has none, `confirm` writes
`from: interview` onto the line, and `freeze` notices the unsourced ids at a rung floor without
ever refusing for them. The tail is Must TEXT, so the direction digest moves with it.
SIX consecutive T2 fresh-session reads on this one task; five FOUND a real defect and every one of
them was the same shape — ONE FACT, TWO READERS. Each fix narrowed the gap a level: template → raw
body → fence view → section view → multiplicity → the SEAL → the WRITER's own regex.
· The SEAL read a fourth view of `## RULES`. `direction_digest` and `binding_digest` sliced the
RAW body with `_section_of`, which is fence-blind, indent-blind and matches the heading only
when spelled exactly `## RULES`. On a node carrying a fenced `## RULES` example, or a heading
spelled `## RULES (frozen)` or `## RULES:` or indented, the whole Must/Reject payload sat
OUTSIDE the seal: `confirm` wrote a Must's tail and the digest did not move (R:DIGESTDRIFT
verbatim), and a frozen Must could be replaced under a build with NO drift refusal at all.
· The WRITER kept a third pattern for what a Must line is — looser on the left than `RULE_ID`,
stricter on the right. `- M2:` was a Must to the gate and to the question and invisible to the
writer, so `interview --answer M2=confirm` exited 0 and wrote nothing: the node then carried an
`act: interview` stamp attesting the answer beside a freeze notice saying the Must had no
source, forever, with its own `next:` verb unactionable.
Two named readers now, and everything asks them:
`authored_section(body, heading)` — where an AUTHORED section is. `direction_digest`,
`binding_digest`, `_open_decisions`, `placeholders_in` and `rules_of` route through it.
`must_lines(body) -> {index: id}` — which LINE is a Must. The question, the freeze notice and the
interview's writer read it, so a view of WHERE and a grammar for WHAT live in one place.
Measured before routing anything: RULES/CHECKS/EDGES/ASSUMPTIONS move 0 of 155 live nodes, so no
frozen node is stranded. `## EXIT` moves 26 milestones and was deliberately left unrouted (M58).
The check that ends the family is DIFFERENTIAL — it asserts every reader of the fact returns the
same answer, and it asks the engine's own `must_lines` rather than re-stating the pattern. The
first draft of it re-implemented the pattern and passed on nothing; that was caught and repointed.
observes-slot — `## PLAN` carries `- O<n> covers: … · signal … · window … · threshold … · action
alert|rollback`. `freeze` notices a human-floor node with no observes, `doctor` names a malformed
one, and `brief` renders `<observes>`.
refute-tier-floor — verify.md's ladder said a T1 read is "a prelude, never the rung's answer" and
the gate had never read it, so a payment task's own builder could refute its own green with
`--tier T1` and pass a human floor. `_refute_of` now returns `(outcome, tier)` from the SAME stamp;
at computed floor `human` a claim outside `SIGNING_TIERS` is refused (R:SELFREFUTE), at `plan` it is
a notice on the success line. The tier stays a CLAIM — read from `outcome:` and `tier:` and nothing
else, never `by:`, `probes:`, `changed:` or who signs the gate (R:TIERJUDGED). The floor is the
COMPUTED one: a `sensitivity: data` task lifted to `human` by `sensitive_paths:` is refused too,
which a mutant swapping the computed floor for the declared one exposed as unbound.
The rung is an ALLOWLIST, not a denylist, and that came from the security lens the human floor
demands. A denylist sent every value it had not enumerated to the permissive branch, so
`tier: "T1 "` — one trailing space, inside quotes — rendered in the engine's OWN `## EVIDENCE` view
as `tier T1`, drew nothing from `doctor`, and recorded a human-floor PASS: a well-formed stamp
attesting nothing. `SIGNING_TIERS` is DERIVED as `REFUTE_TIERS` minus T1, so it follows if the
ladder moves, and `sensitivity_floor`'s R:SILENT_FLOOR law — an unreadable declaration is one the
engine cannot honour, so it floors UP — now governs the control as well as the declaration.
The refusal states the reason TRUE of the value it read (`T1` is a prelude; `T4` is no tier the
ladder recognises), because M1 never quotes its refusal. The plan-floor NOTICE is M2's frozen
literal, because M2 does — a Must that states a string IS that string, and threading the new
clause through it drifted off the contract invisibly until a read caught it.
residue-by-kind (direct) — verify.md froze three residue lenses for every task alike, so a `ui`
task was asked about concurrency and an `infra` one was never asked for its rollback path. A fourth
lens is owed BY KIND, enumerated over `PERSONA_TASK_KINDS`, and intake.md's direct row says the
router owes it too. Funded by compression: the skill surface is line-neutral vs HEAD.
One frozen-contract collision, resolved by the human at the seam rather than by weakening a test.
`refute-tier-and-changed` (done, in the closed milestone `refute-at-t2`) froze R:JUDGED with a
STRUCTURAL proxy: the string `tier` must not appear in `_refute_of`. The successor makes the gate
read that key, and reading a key the author wrote is not judging it — the proxy could no longer
tell the two apart. Both of its checks now bind the property itself, parsed rather than grepped:
the rung reads exactly {act, receipt, outcome, tier} and `changed:` is read by nothing. Stronger
than string-absence; R:JUDGED's claim intact.
Reported, not fixed — each filed as a lesson for the milestone seam:
M58 `_open_decisions`' `## EXIT` loop still uses the raw slicer (26 live milestones would be
re-interviewed — a migration, not a task-scoped fix).
M59 an unreadable `tier:` fails OPEN at a human floor (`t2`, `T4`, `0` all pass; only exact `T1`
or absent refuses), against `sensitivity_floor`'s floor-UP law in the same file. Closing it
widens frozen M1, so it needs a human refreeze.
M60 never mutation-probe the live tree, and never hash it to re-aim a pin: ENGINE_MD5 was aimed
at 3e6e3902… computed off a tree carrying a concurrent read's mutant, so the pin certified a
mutation nobody reviewed. Corrected to 70d1db24…
M61 the differential check compares `must_lines().values()` (ids) while the writer consumes
`.keys()` (indices), so index fidelity is proven by fuzz and bound by nothing.
Every check was proven red first and mutation-proved: 5 mutants on the seal (one per routed site),
5 on the readers, 7 on the tier rung, 3 on the re-aimed R:JUDGED guard — all killed. Two rules
(E9, E12) turned out to be bound only by PARAMETRIZED checks, which bind nothing at the gate
(pytest reports `name[param]`, the node's CHECKS name `name`); both are now in-body sweeps over the
same tables, re-proved against the same mutants. Four add.py twins mirrored, three engine pins
re-aimed by hand with the task and the reason, three skill trees identical.
author: Tin Dang
…man's signature stands it down
`add learn --evidence <sha>` now reads what that commit touched and refuses
R:QUICKSIZEUP when any path matches an `index.md` `sensitive_paths:` pattern and
no node has routed it. A change under the security floor is a node, however
small — the direct lane may not absorb it.
The floor is checked FIRST and always wins, and it reads EVERY lesson, not only
a `quick:` one: a floor a prefix can turn off is not a floor. The owner half —
a path an open frozen Task's `scope:` already owns — stays on `quick:`, because
`learn` takes a sha OR a task cid and never both, so a lesson about a Task's own
work cites that Task's own commit. Each half names a `next:` the author can
actually take, and no refusal advertises dropping the prefix as a way past.
What "routed" means took seven T2 refutes to get right, and each of the first
six fixes was correct about the defect it named while the same class sat one
gate to the left. The class is ONE FACT, MANY READERS: "does this scope entry
cover this path" had six readers through three matchers.
- `_scope_files`/`_scope_holds`/`_scope_candidates` are lifted out of
`scope_digest`, so the freshness set and every other question about an entry
give the SAME answer. Read with `fnmatch`, `--scope '**'` disarmed every
sensitive path in the bundle while the node held no files at all.
- Routing requires the node be frozen AT HUMAN AUTHORITY. `authority_for` reads
an entry with `_paths_touch`, where `_paths_touch('**/*','src/auth/**')` is
False — so the interview never armed and a bare freeze stamped `process`,
while the freshness set says `'**/*'` holds every file. The one shape
invisible to the human-authority gate was exactly the shape that holds
everything. The two readers still disagree; the disagreement now fails CLOSED.
- And the freeze must be one a PERSON signed. `freeze` writes `authority: human`
whenever the floor it computes is human, the default `--by` is `cli`, and
`interview_gap` has nothing to ask when the author left ASSUMPTIONS empty — so
`add new Persona p --scope src/auth/token.py` then `add freeze p` took the
floor down in two commands with nobody anywhere. Reading the computed floor
back asks the engine whether the ENGINE thought a human was owed, never
whether one signed.
`_in_bundle_frame` is the one reader of where the bundle sits: `scope:` and
`sensitive_paths:` are written relative to the bundle PARENT while `diff-tree`
prints repo-root-relative paths, and both walkers had that fact separately and
wrong in both. `--show-prefix` is read unstripped, because a directory is
entitled to a leading space and the default strip ate it, silently making the
floor inert for that bundle. A merge is read first-parent, a root commit only
when git is not shallow, and `-z` keeps a path git would otherwise quote.
The lane is never blocked on git (R:LANEBLOCKED): a receipt cid, an unknown sha,
a missing binary and no repo all land, and an entry no walker can resolve names
no file rather than raising — `learn` now reads every node's scope, so one
`/etc/*` anywhere turned every lesson into a traceback. Two read-only git verbs
throughout (R:OUTWARD), spied on all four commit shapes.
`locate` now answers for the floor as well as the owner, so the page the direct
lane reads before an edit says the same thing the refusal will.
25 checks, each proved red first. Five mutations that shipped the file green are
now bound, including three checks that passed on a floor they never reached.
Open at the seam, not closed here: the node's `## CHECKS` declares 23 of 25 and
`gate:` is unstamped — both need the human. Raised for Direction, engine-wide
and outside this contract: `freeze` should refuse the default `--by cli` at a
computed human floor, and should refuse non-lifecycle types.
author: Tin Dang
…k lane stops at a HARD-STOP Closes out `loop-that-closes` short of its own exit: ten of eleven EXIT boxes are checked, and box nine stays open because `quick-lane-tripwire` is gated HARD-STOP on a finding an eighth T2 read produced and this session reproduced. THE FINDING. A human freeze stamp routes any path later hand-added to that node's `scope:`. `scope:` is in neither seal — `direction_digest` covers RULES, CHECKS and `gives:`, `binding_digest` covers EDGES and probed A-ids — so `_scoped_by_any` checks that a freeze stamp signed `human:` exists and that A17 computes `human` now, and never recomputes a digest, because there is none to recompute. The signature attests a scope the signer never saw. Reproduced with a negative control that refuses: hand-add one entry with sed — no verb, no new stamp, no name typed — and the identical lesson lands while both stamp digests still verify and `doctor` reports nothing. That is worse than the `--by "human:X"` forgery accepted at the round-seven gate: that one leaves a fresh forged line in the ledger; this one leaves a record every reader calls unchanged. Eight rounds on one node, each finding the same class one gate to the left: matcher, then freeze, then authority, then signature, and now the seal's own coverage. The fix changes `freeze`'s stamp shape engine-wide, so it is raised to Direction rather than built here. Also in this commit: - `quick-lane-tripwire` refrozen at interview 9, `human:Tin Dang`, now declaring all 25 checks. The two it under-declared were exactly the checks that closed round seven's HARD-STOP. - Both explores answered and gated. `holdout-that-holds`: a same-repo CI-only ref cannot hold (every ref is readable — 350 here), an encrypted fixture cannot either (loader and workflow are builder-writable), and the one route that holds needs a secret-gated repo owning checks, loader and runner. A receipt binds PRESENCE only — never that it ran in CI, never that the holdout was unread — so T4 needs a new claim, not a widened tier. - `method-health`: the refusal count is derivable for NONE of the seven codes, because a refusal writes nothing by law 3. Compliance-at-close is derivable at zero bytes. Both explores are RISK-ACCEPTED, not PASS: their own frozen CHECKS say they are judged against FINDINGS rather than by pytest, so no receipt can carry ids that bind them, and emitting a junitxml named F1/F2/F3 would manufacture the binding rather than earn it. - The milestone's CLOSE section carries the four promotion counts. Two are derivable (90 floor receipts, 10 red; consumers flagged 0 — and 0 BY CONSTRUCTION, which must not be read as health) and two are not. No notice is promoted to a refusal in 3.7: the frozen rule says "the count decides" but names no threshold and no direction, so as written it cannot decide. - 20 open deltas drained to zero — 12 promoted into the specs' binding decisions, 8 folded as history. - `seal-what-you-signed` opened with its EXIT criteria and four task stubs, deliberately UNFROZEN: ratifying a plan is a human seam and those five criteria have not been put to anyone. Full suite 1966 passed, 7 skipped. Version stays at 3.6.0 — the 3.7.0 bump is held. author: Tin Dang
…is ratified The human interviewed on `seal-what-you-signed` and confirmed all five EXIT criteria; the milestone is frozen at `human:Tin Dang`. Four rulings were put to them rather than assumed, and each is now written into the milestone as a ratified decision: - an undigested legacy stamp routes NOTHING. Every human-frozen node in every existing bundle stops routing on upgrade and earns one human refreeze. Fail-closed is correct for a security control; the cost is a migration note and a `doctor` line so the refreeze is discoverable rather than mysterious. - `freeze` gets BOTH halves as refusals, not notices: it refuses the default `--by cli` at a computed human floor, and refuses a type with no lifecycle to seal. The Persona half is a breaking change with zero blast radius here. - `git ls-tree` is accepted as a third read-only verb, so quick-lane-tripwire's E6 and R:OUTWARD are refrozen to enumerate it. A sixth refreeze is cheap next to a floor whose answer depends on which branch you are standing on. - `quick-lane-tripwire` MOVES to the successor with its HARD-STOP, so `loop-that-closes` can close on what it actually delivered. HOW THE OLD MILESTONE CLOSED, and a lesson about it. It has eleven criteria and ten are met. To close at 10/11 the unmet box was marked `[~]` rather than `[ ]`, and the engine's box counter does not recognise that marker — so `milestone-done` succeeded and reported "10/10". The `verified:` ledger stayed exact (the check stamps name EXIT:1-8,10 and EXIT:11; 9 appears in neither) and the criterion is written out with why it is unmet, but a reader trusting the count over the text would read this as complete. The human ratified closing at 10 of 11. Nobody ratified making the box invisible to the counter, and that is structurally the same move this milestone spent eight T2 rounds refuting: a gate satisfied by hiding its subject. It is stated in the CLOSE section, recorded as M67 and bound into specs/method: a milestone closes over a criterion it did not meet only by recording that criterion, never by making it invisible to the counter The real fix is a MOVED verdict for a criterion, which the engine does not have — a change request in its own right, and it belongs to whoever picks up `seal-what-you-signed`. Full suite 1966 passed, 7 skipped. Version stays 3.6.0. author: Tin Dang
Six human interview receipts (A1-A6, E1-E7, and the four refusal codes) confirmed on 2026-09-24 for decision-manifest-binds-approval. Committed as part of the final 3.x record; the task itself is not frozen and is archived by the 4.0 skill-only cut rather than finished. author: Tin Dang
Bump the eight version declarations, the ENGINE string in both tracked add.py copies, and the dogfood bundle stamps to 3.7.0; re-aim ENGINE_MD5 for the version-string change only. CHANGELOG [3.7.0] records the loop-that-closes milestone, closed 10 of 11 and shipped as-is. 3.7.0 is the final release of the `add` CLI: ADD 4.0 replaces the Python engine with a single markdown skill. author: Tin Dang
Thirty red-first checks landed with six Tasks that are still in direction (decision-manifest-binds-approval, doctor-sees-a-moved-scope, freeze-refuses-an-unsigned, holds-against-the-commit, prevention-earns-binding, receipt-purpose-binds-claim). 3.7.0 ships without those features and 4.0 retires the engine, so each is marked xfail(strict=True) with that reason: the absence is recorded, and a check that starts passing still fails the suite. The two passing tests in those files stay unmarked. Re-aim the SKILL.md prose pin for the 3.7.0 metadata version line. author: Tin Dang
pilotspacex-byte
marked this pull request as ready for review
September 28, 2026 03:59
Comment on lines
+51
to
+54
| ok, note = add.fold( | ||
| bundle, "method", "escaped defect", bind="owner · policy", | ||
| validation="/tasks/filing.d/runs/1.md", | ||
| ) |
|
|
||
| def _run(work, bundle, cid, *, purpose="bound", floor=False): | ||
| command = [sys.executable, "-c", "pass"] | ||
| return add.run(bundle, cid, command, cwd=work, purpose=purpose, floor=floor) |
Comment on lines
+131
to
+133
| result = add.run(bundle, cid, [sys.executable, "-c", | ||
| f"open({str(marker)!r}, 'w').write('ran')"], | ||
| cwd=work, purpose="holdout") |
| import argparse | ||
| import hashlib | ||
| import json | ||
| import os |
| import json | ||
| import os | ||
| import re | ||
| import shutil |
| value = json.loads(file.read_text(encoding="utf-8")) | ||
| if isinstance(value, dict) and "version" in value: | ||
| return value | ||
| except (OSError, ValueError): |
TinDang97
added a commit
that referenced
this pull request
Sep 29, 2026
…I clone CI checks out shallow with no tags, so the PREVIOUS_VERSION selection found no older tag and test_previous_tag_is_strictly_older failed on every run (py3.10 and py3.12, PR #224 and #226) while passing locally, where the clone has every tag. The test now runs the selection in a throwaway repo holding v3.3.0, v3.4.0, v3.5.0, v3.6.0 and v3.10.0; the last one fails a lexical sort, so the check still proves numeric ordering. author: Tin Dang
Both "Tooling tests" jobs (py3.10, py3.12) failed on PR #224 while the suite passed locally, because both checks read state a CI checkout does not have: - test_previous_tag_is_strictly_older ran the PREVIOUS_VERSION selection against the clone's own tags; actions/checkout is shallow and tagless, so no older tag was found. It now runs in a scratch repo holding v3.3.0..v3.6.0 plus v3.10.0, which a lexical sort would get wrong. - test_persona_skill_mirrors_are_byte_identical compared the source persona templates with .add/tooling/, the repo's own gitignored install. A fresh checkout has none, so the comparison saw an empty set. The live comparison now runs only where that install exists. author: Tin Dang
TinDang97
added a commit
that referenced
this pull request
Sep 30, 2026
* feat(skill)!: distil ADD into one skill the model runs without an engine SKILL.md (174 lines) is now the whole method: orient from .add/ and git, size the work into a lane (Quick, Task, Explore, Milestone), then drive Direction (rules, assumptions, failing checks) sealed by a freeze(<slug>) commit, Build to green inside the seal, and Verify by diffing the sealed files against the freeze, running the checks fresh, reading the residue, refuting, and writing one verdict into the task's EVIDENCE. There is no human approval gate; the session report lists every assumption taken so the human can review after. references/format.md defines the hand-maintained ABF-1 subset (PROJECT, specs, milestones, tasks, personas); references/explore.md the research lane. The 3.x phase guides and helper scripts are removed. The three shipped skill trees stay byte-identical; tests/test_skill_only.py guards the shape (one skill, no engine, no CLI instruction, git as the seal). BREAKING CHANGE: the skill no longer calls the `add` CLI; every 3.x verb is retired. author: Tin Dang * refactor!: remove the engine, the agent roster and their tests Delete add-method/tooling/ (add.py, cli.py, templates, pins), the add-worker/add-advisor roster in every tree, the bundled engine copy, FORMAT.md, and the engine-only scripts. Remove tests/engine, tests/skill (3.x prose pins) and the covers-grammar test: what they held no longer exists. The nine starter personas move out of tooling/templates to add-method/personas/ as plain persona files (type: Persona, title:, sources:) that orient on the task file and git instead of CLI verbs. Version parity drops the ENGINE declaration (eight remain). CI no longer materialises a vendored engine before the suite. BREAKING CHANGE: `add` CLI, `.add/tooling/`, and the agent roster are gone. author: Tin Dang * chore(add): archive the 3.x bundle and dogfood a 4.0 one Move this repo's 3.x bundle (145 tasks, 35 milestones, specs, personas) to archive/add-3x-bundle/ beside the 2.x record, and seed a 4.0 bundle: PROJECT.md with four invariants and the test command, five fresh specs whose decisions cite the 2026-09-28 ratification and the benchmark record, the add-4-skill-only milestone, and three personas (method-steward and feature-builder carried and restated for 4.0, plus the security-reviewer starter). engine-notary and gate-security-reviewer guarded engine surfaces that no longer exist and stay archived. The dogfood test now checks the 4.0 shape; .gitignore drops the engine cache entries; SECURITY.md names 4.0.x as the supported line. author: Tin Dang * feat(installer)!: install the skill, personas and a 4.0 bundle — nothing else Both installer twins (bin/cli.js 348 lines, _installer.py 347 + _cli.py) replace ~4k lines and drop the @clack/prompts dependency. A project install refreshes the skill and the vendored persona corpus by stage-then-swap, seeds the starter personas and PROJECT.md without ever overwriting, writes the managed ADD block (new begin marker, legacy one recognised) into CLAUDE.md and AGENTS.md, refreshes a stale 3.x block in .clinerules and similar files only where one exists, and last removes a 3.x .add/tooling/ and ADD's own roster agents when they are recognisably ours. --global installs the skill for the user. --help prints help; unknown flags exit 2 without writing. Packaging ships personas/ and no engine; prepare_bundle.py regenerates _bundled/ for 4.0. publish.yml and teacher-refresh.yml no longer call the engine; CI no longer installs npm deps; marketplace and dependabot text updated. Installer tests rewritten red-green: 89 pass, including a real upgrade over published 3.6.0 artifacts. BREAKING CHANGE: interactive prompts and the removed flags (--force, --no-skill, --stage, prune-data) are gone. author: Tin Dang * docs(book)!: teach ADD 4.0 — files, git and the project's test command Rewrite the book and front-door docs for the skill-only method. Every chapter keeps its idea and swaps the mechanism: the freeze is a freeze(<slug>) commit, the gate is a verdict written into EVIDENCE, the receipt is real command output, and the human reviews after through a report that lists every assumption. The command reference becomes 13 · Files and commits; the bundle chapter condenses format.md; the personas chapters merge; a new chapter 20 maps every 3.x mechanism to its 4.0 equivalent, gives upgrade steps, and says plainly what is given up (mechanical refusal). Appendix D is a real run whose first refute caught a vacuous concurrency check and shows the refreeze. CLAUDE.md, AGENTS.md and .clinerules carry one 4.0 block under the new installer marker. The 3.x book tests and the engine-driven fixture are replaced by tests/book/test_book.py (nav, links, no retired verbs outside the migration page, format parity, worked-example loop) and test_beyond_code.py; shipped-docs, claim-truth and orientation tests are updated. Agent-support claims now match what the installer writes (AGENTS.md and CLAUDE.md). mkdocs build --strict passes. author: Tin Dang * chore(release): prepare 4.0.0 Bump every version declaration to 4.0.0 (package.json, package-lock.json, pyproject.toml, __init__.py, plugin.json; the three SKILL.md trees already read 4.0.0). CHANGELOG [4.0.0] records what changed, what was removed, the installer's new surface, and the 3.x migration path. The three manifest descriptions now describe the skill-only method. Close the add-4-skill-only milestone: all five EXIT criteria ticked with their evidence. Full suite: 128 passed; mkdocs build --strict passes. Publishing stays with the human: merge after the 3.7.0 PR, then tag. author: Tin Dang * docs(blog): animated ADD 3.7 vs 4.0 comparison page A single-file page (React + Framer Motion from esm.sh, no build step) showing what changed between 3.7 and 4.0: the layer stack with and without the engine, one Task's steps and engine calls, the loop stage by stage, the values kept / mechanisms changed / machinery removed, measured size deltas as paired small-multiple bars with a table view, the 3.x benchmark cost, and the trade-off 4.0 makes (mechanical refusal for git-checkable claims). Every figure is measured from the two commits or cited from benchmark/. The two series colors pass the palette validator in light and dark; reduced motion is honored. author: Tin Dang * feat(benchmark): pilot ADD 4.0 against 3.7.0 on the same model Add two arms for a head-to-head pilot. add-4 installs this branch's add-method with the 4.0 installer; add-3x installs the release/3.7.0 worktree, refusing to run unless its HEAD is the pinned fba5445. Both give each workspace its own git repo and baseline commit, so 4.0's freeze/verify commits land in the workspace, not the harness repo, and 3.7's gate has the working tree it requires. The add-skill prompt wrapper is add-loop clause for clause, with the method mechanics swapped (read SKILL.md; freeze/verify commits; verdict in EVIDENCE), and a test holds its length within 15%. add-loop gains the `brief` step 3.7's gate requires. loop_census.py records, from real workspace git state, task files, freeze/refreeze/verify commits, EVIDENCE verdicts, seal integrity and red-before-seal, and reports measured: false rather than a vacuous zero when the workspace cannot be read. The broken `add` arm is retired with a clear refusal; its archived records still score. A test guard fails any test that would reach the real claude CLI. benchmark/tests: 506 passed, 12 skipped. author: Tin Dang * fix(benchmark): census reads colored test output; record the 4.0 vs 3.7 pilot loop_census.red_first missed unittest failures wrapped in ANSI color codes: the codes broke `FAILED (failures=` and the word boundary before `AssertionError`, so a real red run before the seal was not counted. Strip ANSI codes before matching; a test replays the pilot's real output. PILOT-4v3-2026-09-28.md records the first head-to-head (n=1 per cell, same model): quality and loop adherence held for 4.0 on wm1 and amb1, cost was level to 21% higher, total tokens level to 42% lower. Not significant; the next campaign is 3 reps over wm1-3 and amb1. author: Tin Dang * feat(skill): budget turns — inline task template, batched beats The 4.0 pilot put the method's cost in turns, not bytes: every turn re-reads the whole context, and 4.0 ran about 7 turns over a no-method run on wm1 (21 vs 14; 680k vs 420k tokens). The extra turns were a read of references/format.md for the task shape, separate red-run and seal turns, separate evidence writing, and fix-ups. SKILL.md now carries the task template inline, so a Task needs no second file; orients in one command; and adds a Turns section: Direction is two turns (write task + tests together; one command runs the checks and commits the freeze), Build batches files, Verify is two turns (one command for seal diff + checks + regression; one for EVIDENCE + the verify commit). No step is dropped — only round-trips. 197 lines. Guard tests hold the inline template and the turn rule. author: Tin Dang * feat(skill): stub first so the red is for the right reason; census follows the batched loop The turn-budget rerun (2 reps × wm1, amb1) cut 4.0 from 21 to ~15.5 turns and 680k to ~490k tokens on wm1 (no-method floor: 14 / 420k) with quality unchanged and every seal intact. One run in four sealed after a red caused only by an import error; the run that avoided it wrote NotImplementedError stubs first. SKILL.md now says so in the Direction turn rule (198 lines). The census missed three shapes the batched loop produces, each fixed with a test from the real output: a test run inside the seal command now counts toward red-first; the seal is the task file plus the files its CHECKS name, not everything riding in the freeze commit; and a stub's NotImplementedError counts as a right-reason red. PILOT-4v3-2026-09-28.md records the floor, the anatomy and round 2. author: Tin Dang * freeze(close-research-gaps): the 4.0 skill states every closed-loop invariant and the persona-routing roadmap author: Tin Dang * feat(skill): close the closed-loop and persona-routing gaps as stated rules The 4.0 cut dropped the enforcement behind several closed-loop invariants (ADD_3_6_Closed_Loop_Research.md §17) and adopted none of the persona roadmap (ADD_Dynamic_Persona_Research.md). Each gap is now a line on the path the model walks, with no engine and no extra turns on ordinary work: - rules name their source or are marked derived:; checks name a falsifier - a task names its risks:, which pick the persona, a second evidence mode and the residue lenses - floor work: a second reader under the counter-lens before the seal (replacing the removed human pre-approval) and executable counter-lens probes after the build - Verify runs the consumers of a changed gives: surface - Quick work that touches the floor is a Task now - tag only verified work; observes: names what to watch after release - an escape opens a successor (fixes:) that closes only on a bound prevention; zero-yield controls become method deltas Depth lives in two new references read on a trigger: evidence.md (the loop from intent to production) and personas.md (routing, counter-lens, lens: traces, evals, lifecycle). Every starter and repo persona gains covers-risks, evidence and counter-lens. SKILL.md stays at 200 lines. author: Tin Dang * feat(skill): second reader on every floor task; a passed-by security issue leads the report Skill-creator eval, iteration 1 (3 fixtures, new skill vs the pre-change snapshot, claude-sonnet-5) found two regressions in the first cut: - quick-escalation: asked for a typo fix on a log line that writes the user's API token, the new skill filed the leak as an "open risk" under "HARD-STOPs: None"; the old skill led with it as a HARD-STOP. The tripwire only covered the change itself. Now a security issue met in passing stays out of the diff but leads the report as a HARD-STOP. - gives-consumer: a consumed-surface task (floor) skipped the second reader and wrote no lens: line; the invite task skipped the subagent "for this size". The second reader is now every floor task however small: a fresh subagent for security, data and architecture, a cold reread under the counter-lens otherwise; lens: is always written, or "none — why". Iteration 2: every assertion held on all three fixtures (29/29 vs 23/29 for the old skill). The invite run's pre-seal second reader added a rule neither earlier run had: an invite is not consumed if add_member fails. author: Tin Dang * verify(close-research-gaps): PASS Seal intact (freeze 837097b, 0-line diff to head 4e591e8); on the clean tree the guard file passes 13/13 and the full suite 134/134. Six mutation probes on a temp copy were each caught by their guard. The skill-creator eval (3 fixtures, new skill vs the pre-change snapshot) found two regressions in iteration 1, fixed in 4e591e8; iteration 2 held every assertion (29/29 vs 23/29). Extra cost appears only on security-floor work (+46% tokens on invite-expiry), recorded as a method delta to measure. author: Tin Dang * feat(skill): trigger description tuned on a 20-query eval, now covers investigations skill-creator's description loop (claude-opus-5-5, 12 train / 8 held-out queries, 3 runs each; 10 should-trigger, 10 near-miss negatives such as "git add everything", "add a last_login column", writing a lone test file, a PRD, a CI lint fix): - previous description: train 11/12, test 8/8, precision 100%. Its one miss was "investigate why the nightly export slowed down ... cited findings before anyone changes code" — the compressed description had lost the Explore lane. - adopted description: train 12/12, test 8/8, precision 100%. It names evidence-first investigations and says when not to trigger. SKILL.md stays at 200 lines: three sentences that repeated a rule stated elsewhere were folded (the record-real-output line, the floor-only note in Turns, the consumers wrap). author: Tin Dang * test(release): pick the previous tag inside a scratch repo, not the CI clone CI checks out shallow with no tags, so the PREVIOUS_VERSION selection found no older tag and test_previous_tag_is_strictly_older failed on every run (py3.10 and py3.12, PR #224 and #226) while passing locally, where the clone has every tag. The test now runs the selection in a throwaway repo holding v3.3.0, v3.4.0, v3.5.0, v3.6.0 and v3.10.0; the last one fails a lexical sort, so the check still proves numeric ordering. author: Tin Dang * freeze(quality-dimensions): score six code-quality dimensions the oracle cannot see author: Tin Dang * refreeze(quality-dimensions): C4 allows for equivalent mutants; C6 counts 8 annotation slots The fixture's clamp has two mutants no test can kill, so a strong suite tops out at 6/8; C4 now demands a 0.5 gap between strong and weak instead of an unreachable 0.8. C6's expected ratio was an arithmetic slip (3/8, not 3/6). author: Tin Dang * feat(benchmark): six quality dimensions beyond the oracle — edge, mutation, static, security, tests, evidence benchmark/quality.py scores a run's workspace, never writing to it: held-out edge suites for wm1 (19 cases) and amb1 (14), validated against a correct and a sloppy reference app; a seeded AST mutation score of the run's own tests (n/a on a red baseline); AST static quality; security smells; test quality; and, for ADD runs, whether the EVIDENCE regression count matches a fresh rerun. `python -m benchmark.quality <runs…>` writes quality.json per run. author: Tin Dang * refreeze(quality-dimensions): evidence honesty reads the claim shapes real runs write Scoring the 09-29 runs crashed on a `.venv/bin/python` claim (the scoring copy omits .venv) and read n/a on wrapped and unittest-style claims. C10 pins those shapes. author: Tin Dang * fix(benchmark): evidence honesty takes the largest claimed suite count and reruns the whole suite Never replays the agent's own command: it named a .venv the scoring copy omits and crashed the scorer. author: Tin Dang * fix(benchmark): pin the 3.7 arm to the 3.7.0 release head The CI fix on PR #224 (d0af5bb, test-only) moved the release-3-7 worktree off fba5445, so the pin check refused the add-3x arm. The engine code is unchanged; the pin follows the commit that will ship as 3.7.0. author: Tin Dang * verify(quality-dimensions): PASS Seal intact since the last refreeze; check 12/12, benchmark suite 522 passed. Scoring a real run twice is deterministic and leaves the workspace byte-identical. Round-3 results recorded in benchmark/PILOT-4v3-2026-09-29.md. author: Tin Dang * freeze(close-benchmark-gaps): rule coverage, checked guesses, security-only subagents, input robustness author: Tin Dang * feat(skill): every rule gets a check, cheap guesses get checked, subagents only for security From benchmark round 3 and the task-contract audit: two contracts left a Must with no check; found: was almost never used; the second reader cost 2.3x on amb1 (three subagents in one run, a wakeup wait in another); five of nine apps crashed on a null or number body. SKILL.md now asks that every RULES id sit on a covers: line, that cheap guesses be checked now, that input surfaces get a malformed-input check, and caps subagents at one per beat, in the foreground, for security work only. Still 200 lines. author: Tin Dang * refreeze(close-benchmark-gaps): the subagent budget binds every reference, explore.md included A refute probe found references/explore.md still telling the agent to "give parallel subagents disjoint questions", against SKILL.md's one-per-beat, foreground budget. The sealed rule R:CONSISTENT named only evidence.md and personas.md, so its check could not see the conflict. R:CONSISTENT now covers every reference and C2 walks REFERENCES; the check is red until explore.md is brought in line. author: Tin Dang * fix(skill): the Explore lane keeps the subagent budget references/explore.md still told the agent to give parallel subagents disjoint questions, against SKILL.md's one-per-beat, foreground budget. Explore now splits disjoint questions and answers them in turn, with at most one foreground subagent for a question whose reading would flood the context. Mirrored to the bundled and project skill trees. author: Tin Dang * docs(benchmark): round 4 — what ADD 4.0 buys Claude Code and what it only costs Same-day add-4 vs vanilla, wm1 + amb1, n = 3 each, scored on every quality dimension, with each ADD practice mapped to what it moved. The falsifier-per-check practice is the measured value (mutation score +0.17 / +0.26); ASSUMPTIONS change no decision; the malformed-input rule did not transfer (field-level only); Direction is 46% of tokens and about forty turns against the skill's "two". Discloses the 403 reruns, an estimated lost attempt cost, the operator config both arms load, and one escaped timezone defect. author: Tin Dang * verify(close-benchmark-gaps): RISK-ACCEPTED Seal intact against the refreeze 84a0563; the check (15) and the full add-method suite (136) are green on the clean tree, and the three skill trees are identical. Round 4 shows every rule covered in 5 of 6 contracts, found: in every task, and no subagents in 6 of 6 runs. The malformed-input rule was written into every contract but only at field level, so body-level garbage still crashed 4 of 6 apps — accepted and handed to review as a proposed "shape" sweep. A refute probe found explore.md contradicting the subagent budget; refrozen and fixed. author: Tin Dang * fix(benchmark): count Direction in API messages, not transcript lines The round-4 cost ledger counted a message's thinking, text and each tool call as separate turns, reporting Direction at 37–47 turns and 46% of tokens. Deduplicated by message id it is about 20 messages and 41–44%, with Verify at 2–5%. A traced run shows where Direction's messages go: environment probing one command at a time, batched writes, and a fumbled seal commit. The method delta carries the corrected figure. author: Tin Dang * freeze(value-over-ceremony): a short Direction, checks that send real callers' inputs, guesses that fail safe Contract and red guard tests for the round-4 findings: Direction as a batched plan (ground and find the test command in one command, parallel writes in one message, red and seal in one); checks that send the body itself malformed and every allowed value form; a least-privilege reading for silent authorization; ASSUMPTIONS reported costliest-if-wrong first. Round 5, same day, is the behavioural evidence. author: Tin Dang * feat(skill): a three-turn Direction, checks that send real callers' inputs, guesses that fail safe Direction is now a batched plan: one command grounds and finds how the tests run (installing what is missing), one message writes the task file, tests and stubs as parallel writes, and one command runs red and seals. Checks send inputs the way a real caller does: every value form the spec allows (a timestamp with and without an offset) and a malformed body as well as malformed fields. A silence about who may act or see takes the least-privilege reading, and the report lists the assumptions costliest-if-wrong first. The red-first sentence joined the Seal and the residue lenses point at evidence.md, keeping SKILL.md at 200 lines. Mirrored to all three trees. author: Tin Dang * verify(value-over-ceremony): RISK-ACCEPTED Seal intact against b4dbf6d; the check (18) and the full add-method suite (139) are green and the three skill trees are identical. Round 5, same day, shows the input-shape rule changed behaviour: no garbage-body crash in 6 of 6 runs (was 4 of 6), offset-aware timestamps in most tests, every wm1 edge case held. The least-privilege reading, the costliest-first report and the three-turn Direction were stated but mostly not followed: prose that advises moves the model less than an artifact it must write. Vanilla wrote no tests in 2 of 6 runs and halted on the contradictory spec in 1 of 3; ADD delivered and tested in 6 of 6. author: Tin Dang --------- Co-authored-by: Tin Dang <tindang.ht97@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 3.7.0 — the last engine release
Merges
feat/loop-that-closesas-is and cuts 3.7.0. ADD 4.0 (stacked PR to follow) replaces the Python engine with a single markdown skill, so 3.7.0 is the final release of theaddCLI.What ships: the
loop-that-closesmilestone (10 of 11 exit criteria; seeadd-method/CHANGELOG.md[3.7.0]).Deliberately not finished: open 3.x milestones (
seal-what-you-signed,state-that-tells-truth, …) and the quick-lane-tripwire HARD-STOP. They are archived by 4.0, not built.Release commits on top:
f85d381ccommits the recorded decision-manifest interview answers06aafe23bumps all 8 version declarations, theENGINEstring, the dogfood stamps, andENGINE_MD5to 3.7.0fba54456marks 30 red-first checks for six unbuilt direction Tasks asxfail(strict=True), and re-aims the SKILL.md prose pinLocal suite: 2,052 passed · 30 xfailed · 8 skipped. 2 persona tests fail only in a fresh worktree (gitignored
.add/personas-teacher); they pass in the main checkout, and CI runsinitfirst.After merge (human-owned):
Then watch Actions ▸ publish.
author: Tin Dang