Repository navigation
Evidence over tests — acceptance-first checks, an evidence router, and a refute the gate reads - #223
Merged
Merged
Conversation
…asks scaffolded Milestone `evidence-over-tests`, frozen at plan authority against the ratified design memo (ADD 3.7). Seven exit criteria: four are direct-lane changes to the method prose (checks bind not count · acceptance-first checks · evidence router and recipes · verify tiers and probes), three are task nodes (refute-verb · refute-gate-rung · dogfood-and-measure). Lens: method-steward. Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
The gate enforces ≥1 passing check per Must, Reject, filled edge and probed assumption, and one `covers:` may name several referents (FORMAT §8.3). The prose said "one check per Must and per Reject" — a quota an agent reads as N rules → N tests, test count standing in for evidence. Every shipped tree (skill, its two twins, docs 03/12, appendix c/e) now states the rule the gate runs, and says what a check is FOR: to fail on the most plausible wrong implementation. One check may cover several referents when it discriminates each; at a plan or human floor a Must carries two checks of different mode. Guard: tests/skill/test_checks_bind_not_count.py (red first, proven to match the old sentence). SKILL.md re-pinned; the 4 bytes over budget were funded by trimming the terms.md pointer, not by raising the ceiling. Milestone: evidence-over-tests C1. Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
… only bound checks are frozen For code the default frozen check is now an acceptance check: business-readable, run through the port with deterministic adapters behind it, and bound to a filled E-edge written as Given · When · Then — the example a non-technical owner can read and confirm. The mode word after `covers:` is a closed list stated once in the guide and once in the book (`acceptance · property · contract · static · unit · e2e · manual`); the engine never parses it. "The two modes never mix within one task" is gone. Red is narrowed: it proves an absent behavior, never the reading of the rule. PLAN offers an optional, unbound `port:` line. Build's first line freezes only BOUND checks — every other test is the builder's to write and delete, and red-green-refactor is a technique, never what the gate asks. Funded: the ASSUMPTIONS matrix, guess-visibility and micro-spike paragraphs compressed (surface 1495 -> 1495 versus HEAD). Guard: tests/skill/test_acceptance_first_checks.py, red first (7 of 7). Milestone: evidence-over-tests C2. Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
A router keyed on change kind × the engine's COMPUTED floor says which evidence modes a change earns — acceptance, property, contract, static, mutation on changed code, e2e — preferred, not enforced; a divergence is recorded on the PLAN line. Stated once in the guide (direction.md) and once in the book (docs 03) with the same row keys. Property, contract and mutation checks are the same shape as the reconciliation checker: a script that emits JUnit whose threshold is a frozen Must. They ship as runnable scripts beside domains.md (skill/add/scripts/*_check.py) rather than as fenced prose, so the surface budget pays for lines an agent loads, not for code it copies; the guard lifts each SHIPPED script into a real bundle and asserts the receipt reaches `kind: test-ids` with the id bound — and withholds the subject once (a blind suite kills no mutant and the recipe reports it). Funded: receipt prose in verify.md, §1/§3 prose in domains.md, the brief and contract-edge sections in direction.md (surface 1492 -> 1492 versus HEAD). Observe now names regression mining: a production failure becomes a filled edge with a bound acceptance check on the next node. Guard: tests/skill/test_evidence_router_and_recipes.py (red first, 7 of 7). Milestone: evidence-over-tests C3. Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
…an order `add refute <slug> --by <name> (--held | --found "<input>") [--probes N]` appends one `act: refute` stamp naming the receipt it read, the outcome and the probe count. NO-EXEC: it writes no verdict, moves no status, lowers no floor — a lens on a green exactly as `advise` is a lens on a beat. Refusals: R:UNSEALED (no approved intent to read against), R:NORECEIPT (nothing has run), R:NOFINDING (a refutation names its input; one outcome only), R:NOTATASK. A probe count is a count: argparse refuses a negative one (found by this task's own refute-read). `add interview` now compiles every FILLED E-edge as a question (dim `edge`), between the assumptions and the Rejects — the Given/When/Then example is a claim the AI wrote in the human's name, so the human confirms it before the freeze. The scaffold placeholder line is not a question; the digest covers the edge text, so rewording an example re-opens the pass as it does for an assumption. Registries: 28th verb — cli.py, WIRED, four count pins, both READMEs, docs 13, SKILL.md cookbook (funded: 13254/13258 bytes, 175/176 lines), FORMAT §8.4 (the stamp shape and what it does NOT bind). Engine pins re-aimed; four add.py twins and both pin twins synced; message-digest pin re-aimed (98 -> 105, none reworded). Task: refute-verb — frozen at plan (architecture), receipt runs/2.md, refute held (3 probes), gate PASS. Full suite 1500 passed before the receipt. Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
…s a refute citing the gated run `add gate <slug> PASS` now refuses R:UNREFUTED when no `act: refute` stamp cites the gated receipt, and R:REFUTED when the latest citing stamp reads `outcome: refuted`. Both are evidence-class, like `unbriefed`: RISK-ACCEPTED and HARD-STOP are never refused by them. Bound at standard|deep depth on a Task whose COMPUTED floor is plan or human; quick depth, the process floor and the explore lane are exempt. The rung reads presence and outcome only — never the probe count, the note, or who signed (law 3). Chronology is the citation: a refute names the receipt it read, so a refute of an earlier green enters nothing for a later one, and a stamp with no `receipt:` cites nothing. The verify-beat hint (`status` · `todo` · the run receipt's own next line) names `add refute` before the gate on a rung-bound task, so the first contact with the rung is a hint, not a refusal. FORMAT §8.4 states the rung and the exemptions; verify.md §3, gate.md and docs 05 each name it in one line. Three existing gate fixtures now refute before they PASS (security lens, path floor, architecture) — the rung firing on them is the rung working. Pinned refusal set re-aimed to include `unrefuted`. Task: refute-gate-rung — frozen at plan (architecture), receipt runs/2.md, refute held (3 probes), gate PASS through the rung it ships. Full suite 1510 passed before the receipt. Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
…nce, parity-pinned A probe is a check the builder never saw as a target. What a refuter may derive one from is a CLOSED list — instantiate a frozen rule with unused values, compose two frozen rules, walk an implied boundary, vary a swept dimension — and the one thing it may never do is invent a requirement: an expected answer not derivable from frozen RULES, EDGES and interviewed ASSUMPTIONS is a spec silence, a change-request back to Direction, never a finding. A probe that finds a defect graduates into a filled edge at the refreeze; one that holds stays unbound. The list is read by two parties who must agree — the worker at Verify and the advisor in refute mode — so it is pinned byte-identical in verify.md and add-advisor.md (three agent trees). The advisor's refute mode now reads the frozen node BEFORE the diff and ends its Return with the exact `add refute` line. The tier ladder (T0 nobody · T1 same session · T2 fresh session, what the rung asks for at floor ≥ plan · T3 human · T4 protected holdout) is stated in verify.md and docs 05, and admits that T4 is a CI recipe: a prompt is not isolation. Funded: receipt, PASS, residue and freeze prose compressed (surface 1491 -> 1491). Guard: tests/skill/test_verify_tiers_and_probes.py, red first (4 of 5; the twin parity check is green by construction and stays as a drift tripwire). Milestone: evidence-over-tests C6, plus the dogfood-and-measure explore (T7): findings F1–F4 read from this bundle's own nodes and stamps; gate PASS on sources. Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
… drained, no bench yet Four direct-lane commits and three loop-lane tasks. The dogfood explore read the milestone's own nodes and stamps: 0.82 bound checks per referent on both engine tasks, 0 of 10 plan-floor Musts carrying two modes (the rule C1 states was not followed by the session that wrote it), 2 refutes both at tier T1, 6 probes of which 1 changed the build under a `held` outcome. Decision: no bench yet — the memo's trigger counted `--found` and undercounted probes; the next real milestone refutes at T2 (a fresh session per task) and counts probes that changed something. Lessons folded into system (3, bound) and method (3, two bound). Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
| """ | ||
| import json, sys, xml.etree.ElementTree as ET; sys.path.insert(0, ".") # the consumer's pact, verified against the provider | ||
| from api import handle # provider entry: handle(method, path, body) -> (status, body) | ||
| pact = json.load(open("pact.json")); suite = ET.Element("testsuite", name="checks.pact"); bad = 0 |
| """ | ||
| import json, sys, xml.etree.ElementTree as ET; sys.path.insert(0, ".") # the consumer's pact, verified against the provider | ||
| from api import handle # provider entry: handle(method, path, body) -> (status, body) | ||
| pact = json.load(open("pact.json")); suite = ET.Element("testsuite", name="checks.pact"); bad = 0 |
| """ | ||
| import json, sys, xml.etree.ElementTree as ET; sys.path.insert(0, ".") # the consumer's pact, verified against the provider | ||
| from api import handle # provider entry: handle(method, path, body) -> (status, body) | ||
| pact = json.load(open("pact.json")); suite = ET.Element("testsuite", name="checks.pact"); bad = 0 |
| changed = subprocess.run(["git", "diff", "--name-only", "HEAD~1", "--", scope], capture_output=True, text=True).stdout.split() | ||
| SWAPS, killed, total = [(" + ", " - "), ("<=", "<"), ("==", "!="), (" and ", " or ")], 0, 0 | ||
| for path in changed: | ||
| src = open(path).read() |
| src = open(path).read() | ||
| for a, b in SWAPS: | ||
| for m in re.finditer(re.escape(a), src): | ||
| open(path, "w").write(src[:m.start()] + b + src[m.end():]); total += 1 |
| for a, b in SWAPS: | ||
| for m in re.finditer(re.escape(a), src): | ||
| open(path, "w").write(src[:m.start()] + b + src[m.end():]); total += 1 | ||
| killed += subprocess.run(cmd, capture_output=True).returncode != 0; open(path, "w").write(src) |
| changed = subprocess.run(["git", "diff", "--name-only", "HEAD~1", "--", scope], capture_output=True, text=True).stdout.split() | ||
| SWAPS, killed, total = [(" + ", " - "), ("<=", "<"), ("==", "!="), (" and ", " or ")], 0, 0 | ||
| for path in changed: | ||
| src = open(path).read() |
| changed = subprocess.run(["git", "diff", "--name-only", "HEAD~1", "--", scope], capture_output=True, text=True).stdout.split() | ||
| SWAPS, killed, total = [(" + ", " - "), ("<=", "<"), ("==", "!="), (" and ", " or ")], 0, 0 | ||
| for path in changed: | ||
| src = open(path).read() |
…efute, gate, close, and `doctor --sync` backfill FORMAT §5 promised "EVIDENCE receipt / gate · LESSONS harvested at done" and the placeholder guard skipped both sections because "the run and the close fill them". Nothing did: 88 of 113 done tasks carried `receipt: <runs/<n>.md>` beside a `verified[]` that named the real receipt, and 64 carried the `<lesson>` scaffold beside specs full of deltas citing them. The sections become VIEWS of the record, never a second store: - `render_evidence` writes the keyed lines — `receipt:` at `add run`, `refute:` at `add refute`, `gate:` at `add gate` (both paths) — replacing only its own keys; a line the engine did not produce survives byte-for-byte (R:TWOHOMES). - `harvest_lessons` at `done` lists every delta whose evidence cites the task (`/tasks/<slug>.md` or its `.d/` sidecar); with none it writes `none filed` and never borrows a lesson that cites something else (R:MANUFACTURED). - `doctor --sync` recomputes both views for the backlog; `doctor` reports a done node still carrying either scaffold as `evidence_scaffold` (info). - `add learn` tells a caller whose evidence names no task how to be harvested. - FORMAT §5, docs 12 and the guard's comment now say who writes the sections. Live bundle backfilled: 113 tasks' EVIDENCE and LESSONS recomputed, 39 carry harvested lessons, 0 done tasks still scaffolded, second sync is a no-op. Task /tasks/evidence-and-lessons-are-views.md: 12 checks red→green, refuted held (4 probes, T1), gate PASS at plan; full suite 1527 passed. Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo author: Tin Dang
| src = open(path).read() | ||
| for a, b in SWAPS: | ||
| for m in re.finditer(re.escape(a), src): | ||
| open(path, "w").write(src[:m.start()] + b + src[m.end():]); total += 1 |
| for a, b in SWAPS: | ||
| for m in re.finditer(re.escape(a), src): | ||
| open(path, "w").write(src[:m.start()] + b + src[m.end():]); total += 1 | ||
| killed += subprocess.run(cmd, capture_output=True).returncode != 0; open(path, "w").write(src) |
…and probe yield readable from the stamp dogfood-and-measure found both refutes on evidence-over-tests were the builder's own read (T1) with the tier written in free text inside --by, and the one probe that changed the build was stamped `held` because no frozen rule forbade it — so the memo's --found trigger read 0 where the honest count was 1 of 6. - `add refute --tier T1|T2|T3` records the CLAIM of who read the green; T0 is nobody and T4 a CI recipe, both refused (R:BADTIER). Absent when not given — a default would invent a claim nobody made. - `add refute --changed "<what>"` records what the probes moved in the build or the spec while the outcome still held — the bench trigger reads this key. Combines with either outcome. - EVIDENCE `refute:` line renders both; FORMAT §8.4 states them and that `_refute_of` reads neither (notary law 3). - The verify hint and the R:UNREFUTED next line name `--tier T2` (a fresh session). Docs 05/13 and the cookbook line ×3 carry the flags, funded by trimming three cookbook comments (SKILL.md 13254 B, pin held). - This task's own refute ran at T2 (a fresh add-advisor session): 3 probes held; P1 found the library stripping a padded --tier that argparse refuses — membership now on the raw value, the padded case joined the bound check, recorded as `changed:`. Engine twins ×4 synced; ENGINE_MD5 / ENGINE_PKG_MD5 re-aimed; refusal message count 105→106 (+BADTIER). Pre-existing red at HEAD untouched: test_live_bundle_backfill_count (its own precondition retired itself). Task: .add/tasks/refute-tier-and-changed.md · milestone refute-at-t2 author: Tin Dang Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
… T2 is the default, T1 a prelude dogfood-and-measure F2: with the tier ladder stated as a list of options, both refutes on the milestone that shipped it were the builder's own read. The engine now records `--tier`; this makes the prose choose. - verify.md "Who refutes": T2 is the DEFAULT at floor ≥ plan — SPAWN `add-advisor` in refute mode, briefed from the frozen node before the diff, record its line with `--tier T2`; T1 is a prelude, never the rung's answer. The refute-read template now carries `--tier T2` (the T2 refuter of this task found the copy-paste path recorded an uncountable stamp). T0/T4 wording unchanged. - SKILL.md step 3 VERIFY names the spawn before the gate, funded by trimming three cookbook comments (13256 B under the 13258 pin, 175 lines, no budget literal moved; PROSE_PINS re-aimed). - add-worker.md §4 verify bullet: spawn, record `--tier T2`, own read is T1. add-advisor.md's refute Return line carries `--tier T2`. - docs/05 tier table: T2 "the default at floor ≥ plan — spawned, never optional"; T1 "a prelude". - Refuted at T2 by a fresh add-advisor session: 3 probes held, 1 change acted on (recorded as `changed:`). The HEAD-vs-worktree checks in the new skill test declare themselves TRIPWIREs (R:SILENTTRIPWIRE). Task: .add/tasks/t2-refute-default.md · milestone refute-at-t2 author: Tin Dang Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
…ce mode — a notice, never a refusal dogfood-and-measure F1: the session that wrote "at a plan floor a Must carries two checks of different mode" froze 0 of 10 such Musts one task later. Prose nobody follows binds nothing; a refusal would grow ceremony back on the mechanical lane. So the freeze names them and judges nothing. - `_check_lines` / `check_modes` / `single_mode_musts`: the mode word is the first token after the second `·` of a CHECKS line, from the closed set acceptance · property · contract · static · unit · e2e · manual, stripped of backticks/parentheses. Keyed by LINE, not by id. - `freeze` at the refute rung's arming (standard|deep × floor plan|human × not explore) appends `notice: M1 (acceptance), M3 (contract) carry one evidence mode — a plan-floor Must carries two (direction.md § router)` to its success note; the stamp is written first. `add todo` carries the count. Quick · process floor · explore · Rejects · edges print nothing. - FORMAT §6.2 states it; docs/03 and direction.md ×3 name it (line-neutral). - Refuted at T2 three times by fresh add-advisor sessions: refute 1 FOUND a duplicated check id letting the last line's mode overwrite the first (fixed, graduated into edge E3, refrozen); refute 2 FOUND FORMAT §6.2 documenting a notice shape the engine never emits (fixed, the prose check now matches the engine's format string); refute 3 recorded below. - Direct fix (milestone exit 4): docs/03 no longer says acceptance mode is for "where no unit test fits" — every screen state is a filled edge with a bound acceptance check. Engine twins ×4 synced; ENGINE_MD5 re-aimed (comment repaired to one pointer). Pre-existing red at HEAD untouched: test_live_bundle_backfill_count. Task: .add/tasks/two-mode-notice.md · milestone refute-at-t2 author: Tin Dang Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
…o deltas folded Every refute on this milestone was recorded at tier T2 by a spawned add-advisor session: 6 refutes, 3 refuted (each graduated into an edge and refrozen), 3 held; 5 of 18 probes changed the build or the spec. The bench trigger does not fire. The single-mode notice fired on 3 of 3 freezes — the two-mode rule's first measured compliance is ~10%; the promote-or-drop decision waits for n>3. author: Tin Dang Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
…nstead of reading it from the live bundle test_live_bundle_backfill_count asserted the live bundle still carried the pre-3.6 EVIDENCE/LESSONS scaffold, then synced and asserted it was gone. c9b4feb backfilled the live bundle in the same commit, so the precondition died and the tooling jobs on PR #223 went red on it. The check now copies the bundle, re-seeds the scaffold into every done task (117 on this bundle, floor 10), syncs, and asserts none remains. Proven red by withholding the sync: 117 still carry a scaffold. Quick lane · lesson filed under tdd. author: Tin Dang Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
… READMEs say what 3.6 ships The tagline and the trust claims still described 3.5: "trust comes from passing tests". Since evidence-over-tests and refute-at-t2 the gate refuses a PASS nobody tried to refute at a plan-or-human floor, the default frozen check for code is an acceptance check through a port, property/contract/mutation checkers earn the same receipt, and the freeze names a Must still on one kind of evidence. - tagline: "Trust comes from evidence that survived a refute" - 🔬 highlight: checks pass on a fresh receipt AND a session that did not build it tried to break the green; new 🎯 "Evidence, not test count" bullet, registered in the promised-capabilities guard (verb:refute · engine:CHECK_MODES) - vanilla-vs-ADD "what you verify" row; How ADD Works Verify beat; the package README's flow line, non-negotiable 2, and the add-advisor refute-mode sentence Identity lines untouched. Quick lane. author: Tin Dang Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Evidence over tests — trust on independent evidence against frozen intent
Milestone
evidence-over-tests, closed 7/7. Design memo: https://claude.ai/code/artifact/facb60e5-a047-4db3-8edc-90312dc251c7Why. The book said "one check per Must" and "trust comes from passing tests" while the gate enforces ≥1 discriminating check per referent and FORMAT §10 concedes a check can assert nothing. The one act that reads a green against its intent — the refute — was prose in
verify.mdthat nothing recorded. A 2026 research pass (Vaccari ATs · Example Mapping · SpecBench · TDFlow · mutation-vs-coverage) says the same session writing rule, check and code shares one misunderstanding three times; the fix is a readable oracle a human confirms and an independent read the record can show.Direct lane (red first, one commit each)
skill/add/scripts/; the guard runs each through a real bundle and assertskind: test-ids.verify.md⇔add-advisor.md; the tier ladder T0–T4 (T4 is admitted to be a CI recipe).ADD loop (frozen at plan · receipts · refutes · gates)
add refute— 28th verb. Stamps who read the green, against which receipt, held or refuted, with probe count. NO-EXEC, no verdict.add interviewcompiles filled edges as questions. Its own refute-read found--probes -1stamped verbatim; argparse now refuses it.add gate PASSrefusesR:UNREFUTED/R:REFUTEDat standard|deep depth on a plan-or-human floor; evidence-class (RISK-ACCEPTED and HARD-STOP untouched); quick · process · explore exempt. The run receipt's next line is now derived, so the first contact is a hint.Measured on this milestone
Decision: no bench yet. The memo's trigger (
--found== 0) undercounts probes that changed the build under aheldoutcome. The next real milestone refutes at T2 (a freshadd-advisorsession per task) and counts probes that changed something.Not in this PR
Full suite: 1515 passed, 7 skipped.
https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo