Skip to content

Evidence over tests — acceptance-first checks, an evidence router, and a refute the gate reads - #223

Merged
TinDang97 merged 16 commits into
mainfrom
feat/evidence-over-tests
Sep 11, 2026
Merged

TinDang97 merged 16 commits into
mainfrom
feat/evidence-over-tests

Conversation

@TinDang97

Copy link
Copy Markdown
Collaborator

Evidence over tests — trust on independent evidence against frozen intent

Milestone evidence-over-tests, closed 7/7. Design memo: https://claude.ai/code/artifact/facb60e5-a047-4db3-8edc-90312dc251c7

Why. The book said "one check per Must" and "trust comes from passing tests" while the gate enforces ≥1 discriminating check per referent and FORMAT §10 concedes a check can assert nothing. The one act that reads a green against its intent — the refute — was prose in verify.md that nothing recorded. A 2026 research pass (Vaccari ATs · Example Mapping · SpecBench · TDFlow · mutation-vs-coverage) says the same session writing rule, check and code shares one misunderstanding three times; the fix is a readable oracle a human confirms and an independent read the record can show.

Direct lane (red first, one commit each)

  • C1 "one check per Must" is gone from every shipped tree; the rule reads as the gate enforces it — at least one check per referent, written to fail on the plausible wrong implementation; one check may cover several.
  • C2 Acceptance first for code: the readable example is a filled Given/When/Then edge an acceptance check covers; the mode word is a closed list stated once in the guide and once in the book; Build freezes only BOUND checks.
  • C3 The evidence router (change kind × computed floor) and property · contract · mutation-on-changed-code recipes as runnable scripts under skill/add/scripts/; the guard runs each through a real bundle and asserts kind: test-ids.
  • C6 The closed probe derivation list, parity-pinned verify.md ⇔ add-advisor.md; the tier ladder T0–T4 (T4 is admitted to be a CI recipe).

ADD loop (frozen at plan · receipts · refutes · gates)

  • add refute — 28th verb. Stamps who read the green, against which receipt, held or refuted, with probe count. NO-EXEC, no verdict. add interview compiles filled edges as questions. Its own refute-read found --probes -1 stamped verbatim; argparse now refuses it.
  • The rung — add gate PASS refuses R:UNREFUTED / R:REFUTED at standard|deep depth on a plan-or-human floor; evidence-class (RISK-ACCEPTED and HARD-STOP untouched); quick · process · explore exempt. The run receipt's next line is now derived, so the first contact is a hint.
  • dogfood-and-measure — explore, gated on findings read from the bundle.

Measured on this milestone

measure value
bound checks per referent (both engine tasks) 0.82
plan-floor Musts carrying two modes 0 of 10
refutes recorded · tier 2 · both T1
probes run · probes that changed the build 6 · 1
human minutes at any freeze 0

Decision: no bench yet. The memo's trigger (--found == 0) undercounts probes that changed the build under a held outcome. The next real milestone refutes at T2 (a fresh add-advisor session per task) and counts probes that changed something.

Not in this PR

  • Version bump (eight declarations stay 3.6.0) — release work.
  • The README tagline — human-owned wording.

Full suite: 1515 passed, 7 skipped.

https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo

…asks scaffolded

Milestone `evidence-over-tests`, frozen at plan authority against the ratified
design memo (ADD 3.7). Seven exit criteria: four are direct-lane changes to the
method prose (checks bind not count · acceptance-first checks · evidence router
and recipes · verify tiers and probes), three are task nodes (refute-verb ·
refute-gate-rung · dogfood-and-measure). Lens: method-steward.

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
The gate enforces ≥1 passing check per Must, Reject, filled edge and probed
assumption, and one `covers:` may name several referents (FORMAT §8.3). The
prose said "one check per Must and per Reject" — a quota an agent reads as
N rules → N tests, test count standing in for evidence. Every shipped tree
(skill, its two twins, docs 03/12, appendix c/e) now states the rule the gate
runs, and says what a check is FOR: to fail on the most plausible wrong
implementation. One check may cover several referents when it discriminates
each; at a plan or human floor a Must carries two checks of different mode.

Guard: tests/skill/test_checks_bind_not_count.py (red first, proven to match
the old sentence). SKILL.md re-pinned; the 4 bytes over budget were funded by
trimming the terms.md pointer, not by raising the ceiling.

Milestone: evidence-over-tests C1.

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
… only bound checks are frozen

For code the default frozen check is now an acceptance check: business-readable,
run through the port with deterministic adapters behind it, and bound to a filled
E-edge written as Given · When · Then — the example a non-technical owner can read
and confirm. The mode word after `covers:` is a closed list stated once in the guide
and once in the book (`acceptance · property · contract · static · unit · e2e ·
manual`); the engine never parses it. "The two modes never mix within one task" is
gone. Red is narrowed: it proves an absent behavior, never the reading of the rule.
PLAN offers an optional, unbound `port:` line. Build's first line freezes only BOUND
checks — every other test is the builder's to write and delete, and red-green-refactor
is a technique, never what the gate asks.

Funded: the ASSUMPTIONS matrix, guess-visibility and micro-spike paragraphs compressed
(surface 1495 -> 1495 versus HEAD). Guard: tests/skill/test_acceptance_first_checks.py,
red first (7 of 7).

Milestone: evidence-over-tests C2.

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
A router keyed on change kind × the engine's COMPUTED floor says which evidence
modes a change earns — acceptance, property, contract, static, mutation on
changed code, e2e — preferred, not enforced; a divergence is recorded on the
PLAN line. Stated once in the guide (direction.md) and once in the book (docs 03)
with the same row keys.

Property, contract and mutation checks are the same shape as the reconciliation
checker: a script that emits JUnit whose threshold is a frozen Must. They ship as
runnable scripts beside domains.md (skill/add/scripts/*_check.py) rather than as
fenced prose, so the surface budget pays for lines an agent loads, not for code it
copies; the guard lifts each SHIPPED script into a real bundle and asserts the
receipt reaches `kind: test-ids` with the id bound — and withholds the subject
once (a blind suite kills no mutant and the recipe reports it).

Funded: receipt prose in verify.md, §1/§3 prose in domains.md, the brief and
contract-edge sections in direction.md (surface 1492 -> 1492 versus HEAD).
Observe now names regression mining: a production failure becomes a filled edge
with a bound acceptance check on the next node.

Guard: tests/skill/test_evidence_router_and_recipes.py (red first, 7 of 7).
Milestone: evidence-over-tests C3.

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
…an order

`add refute <slug> --by <name> (--held | --found "<input>") [--probes N]`
appends one `act: refute` stamp naming the receipt it read, the outcome and the
probe count. NO-EXEC: it writes no verdict, moves no status, lowers no floor —
a lens on a green exactly as `advise` is a lens on a beat. Refusals: R:UNSEALED
(no approved intent to read against), R:NORECEIPT (nothing has run), R:NOFINDING
(a refutation names its input; one outcome only), R:NOTATASK. A probe count is
a count: argparse refuses a negative one (found by this task's own refute-read).

`add interview` now compiles every FILLED E-edge as a question (dim `edge`),
between the assumptions and the Rejects — the Given/When/Then example is a claim
the AI wrote in the human's name, so the human confirms it before the freeze.
The scaffold placeholder line is not a question; the digest covers the edge text,
so rewording an example re-opens the pass as it does for an assumption.

Registries: 28th verb — cli.py, WIRED, four count pins, both READMEs, docs 13,
SKILL.md cookbook (funded: 13254/13258 bytes, 175/176 lines), FORMAT §8.4 (the
stamp shape and what it does NOT bind). Engine pins re-aimed; four add.py twins
and both pin twins synced; message-digest pin re-aimed (98 -> 105, none reworded).

Task: refute-verb — frozen at plan (architecture), receipt runs/2.md, refute held
(3 probes), gate PASS. Full suite 1500 passed before the receipt.

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
…s a refute citing the gated run

`add gate <slug> PASS` now refuses R:UNREFUTED when no `act: refute` stamp
cites the gated receipt, and R:REFUTED when the latest citing stamp reads
`outcome: refuted`. Both are evidence-class, like `unbriefed`: RISK-ACCEPTED
and HARD-STOP are never refused by them. Bound at standard|deep depth on a Task
whose COMPUTED floor is plan or human; quick depth, the process floor and the
explore lane are exempt. The rung reads presence and outcome only — never the
probe count, the note, or who signed (law 3).

Chronology is the citation: a refute names the receipt it read, so a refute of
an earlier green enters nothing for a later one, and a stamp with no `receipt:`
cites nothing. The verify-beat hint (`status` · `todo` · the run receipt's own
next line) names `add refute` before the gate on a rung-bound task, so the
first contact with the rung is a hint, not a refusal.

FORMAT §8.4 states the rung and the exemptions; verify.md §3, gate.md and docs
05 each name it in one line. Three existing gate fixtures now refute before
they PASS (security lens, path floor, architecture) — the rung firing on them
is the rung working. Pinned refusal set re-aimed to include `unrefuted`.

Task: refute-gate-rung — frozen at plan (architecture), receipt runs/2.md,
refute held (3 probes), gate PASS through the rung it ships. Full suite 1510
passed before the receipt.

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
…nce, parity-pinned

A probe is a check the builder never saw as a target. What a refuter may derive
one from is a CLOSED list — instantiate a frozen rule with unused values, compose
two frozen rules, walk an implied boundary, vary a swept dimension — and the one
thing it may never do is invent a requirement: an expected answer not derivable
from frozen RULES, EDGES and interviewed ASSUMPTIONS is a spec silence, a
change-request back to Direction, never a finding. A probe that finds a defect
graduates into a filled edge at the refreeze; one that holds stays unbound.

The list is read by two parties who must agree — the worker at Verify and the
advisor in refute mode — so it is pinned byte-identical in verify.md and
add-advisor.md (three agent trees). The advisor's refute mode now reads the frozen
node BEFORE the diff and ends its Return with the exact `add refute` line.

The tier ladder (T0 nobody · T1 same session · T2 fresh session, what the rung asks
for at floor ≥ plan · T3 human · T4 protected holdout) is stated in verify.md and
docs 05, and admits that T4 is a CI recipe: a prompt is not isolation.

Funded: receipt, PASS, residue and freeze prose compressed (surface 1491 -> 1491).
Guard: tests/skill/test_verify_tiers_and_probes.py, red first (4 of 5; the twin
parity check is green by construction and stays as a drift tripwire).

Milestone: evidence-over-tests C6, plus the dogfood-and-measure explore (T7):
findings F1–F4 read from this bundle's own nodes and stamps; gate PASS on sources.

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
… drained, no bench yet

Four direct-lane commits and three loop-lane tasks. The dogfood explore read the
milestone's own nodes and stamps: 0.82 bound checks per referent on both engine
tasks, 0 of 10 plan-floor Musts carrying two modes (the rule C1 states was not
followed by the session that wrote it), 2 refutes both at tier T1, 6 probes of
which 1 changed the build under a `held` outcome. Decision: no bench yet — the
memo's trigger counted `--found` and undercounted probes; the next real milestone
refutes at T2 (a fresh session per task) and counts probes that changed something.

Lessons folded into system (3, bound) and method (3, two bound).

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
"""
import json, sys, xml.etree.ElementTree as ET; sys.path.insert(0, ".") # the consumer's pact, verified against the provider
from api import handle # provider entry: handle(method, path, body) -> (status, body)
pact = json.load(open("pact.json")); suite = ET.Element("testsuite", name="checks.pact"); bad = 0
"""
import json, sys, xml.etree.ElementTree as ET; sys.path.insert(0, ".") # the consumer's pact, verified against the provider
from api import handle # provider entry: handle(method, path, body) -> (status, body)
pact = json.load(open("pact.json")); suite = ET.Element("testsuite", name="checks.pact"); bad = 0
"""
import json, sys, xml.etree.ElementTree as ET; sys.path.insert(0, ".") # the consumer's pact, verified against the provider
from api import handle # provider entry: handle(method, path, body) -> (status, body)
pact = json.load(open("pact.json")); suite = ET.Element("testsuite", name="checks.pact"); bad = 0
changed = subprocess.run(["git", "diff", "--name-only", "HEAD~1", "--", scope], capture_output=True, text=True).stdout.split()
SWAPS, killed, total = [(" + ", " - "), ("<=", "<"), ("==", "!="), (" and ", " or ")], 0, 0
for path in changed:
src = open(path).read()
src = open(path).read()
for a, b in SWAPS:
for m in re.finditer(re.escape(a), src):
open(path, "w").write(src[:m.start()] + b + src[m.end():]); total += 1
for a, b in SWAPS:
for m in re.finditer(re.escape(a), src):
open(path, "w").write(src[:m.start()] + b + src[m.end():]); total += 1
killed += subprocess.run(cmd, capture_output=True).returncode != 0; open(path, "w").write(src)
changed = subprocess.run(["git", "diff", "--name-only", "HEAD~1", "--", scope], capture_output=True, text=True).stdout.split()
SWAPS, killed, total = [(" + ", " - "), ("<=", "<"), ("==", "!="), (" and ", " or ")], 0, 0
for path in changed:
src = open(path).read()
changed = subprocess.run(["git", "diff", "--name-only", "HEAD~1", "--", scope], capture_output=True, text=True).stdout.split()
SWAPS, killed, total = [(" + ", " - "), ("<=", "<"), ("==", "!="), (" and ", " or ")], 0, 0
for path in changed:
src = open(path).read()
…efute, gate, close, and `doctor --sync` backfill

FORMAT §5 promised "EVIDENCE receipt / gate · LESSONS harvested at done" and the
placeholder guard skipped both sections because "the run and the close fill them".
Nothing did: 88 of 113 done tasks carried `receipt: <runs/<n>.md>` beside a
`verified[]` that named the real receipt, and 64 carried the `<lesson>` scaffold
beside specs full of deltas citing them.

The sections become VIEWS of the record, never a second store:
- `render_evidence` writes the keyed lines — `receipt:` at `add run`, `refute:` at
  `add refute`, `gate:` at `add gate` (both paths) — replacing only its own keys;
  a line the engine did not produce survives byte-for-byte (R:TWOHOMES).
- `harvest_lessons` at `done` lists every delta whose evidence cites the task
  (`/tasks/<slug>.md` or its `.d/` sidecar); with none it writes `none filed` and
  never borrows a lesson that cites something else (R:MANUFACTURED).
- `doctor --sync` recomputes both views for the backlog; `doctor` reports a done
  node still carrying either scaffold as `evidence_scaffold` (info).
- `add learn` tells a caller whose evidence names no task how to be harvested.
- FORMAT §5, docs 12 and the guard's comment now say who writes the sections.

Live bundle backfilled: 113 tasks' EVIDENCE and LESSONS recomputed, 39 carry
harvested lessons, 0 done tasks still scaffolded, second sync is a no-op.
Task /tasks/evidence-and-lessons-are-views.md: 12 checks red→green, refuted
held (4 probes, T1), gate PASS at plan; full suite 1527 passed.

Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
author: Tin Dang
src = open(path).read()
for a, b in SWAPS:
for m in re.finditer(re.escape(a), src):
open(path, "w").write(src[:m.start()] + b + src[m.end():]); total += 1
for a, b in SWAPS:
for m in re.finditer(re.escape(a), src):
open(path, "w").write(src[:m.start()] + b + src[m.end():]); total += 1
killed += subprocess.run(cmd, capture_output=True).returncode != 0; open(path, "w").write(src)
…and probe yield readable from the stamp

dogfood-and-measure found both refutes on evidence-over-tests were the
builder's own read (T1) with the tier written in free text inside --by,
and the one probe that changed the build was stamped `held` because no
frozen rule forbade it — so the memo's --found trigger read 0 where the
honest count was 1 of 6.

- `add refute --tier T1|T2|T3` records the CLAIM of who read the green;
  T0 is nobody and T4 a CI recipe, both refused (R:BADTIER). Absent when
  not given — a default would invent a claim nobody made.
- `add refute --changed "<what>"` records what the probes moved in the
  build or the spec while the outcome still held — the bench trigger
  reads this key. Combines with either outcome.
- EVIDENCE `refute:` line renders both; FORMAT §8.4 states them and that
  `_refute_of` reads neither (notary law 3).
- The verify hint and the R:UNREFUTED next line name `--tier T2` (a
  fresh session). Docs 05/13 and the cookbook line ×3 carry the flags,
  funded by trimming three cookbook comments (SKILL.md 13254 B, pin held).
- This task's own refute ran at T2 (a fresh add-advisor session): 3
  probes held; P1 found the library stripping a padded --tier that
  argparse refuses — membership now on the raw value, the padded case
  joined the bound check, recorded as `changed:`.

Engine twins ×4 synced; ENGINE_MD5 / ENGINE_PKG_MD5 re-aimed; refusal
message count 105→106 (+BADTIER). Pre-existing red at HEAD untouched:
test_live_bundle_backfill_count (its own precondition retired itself).

Task: .add/tasks/refute-tier-and-changed.md · milestone refute-at-t2
author: Tin Dang
Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
… T2 is the default, T1 a prelude

dogfood-and-measure F2: with the tier ladder stated as a list of options,
both refutes on the milestone that shipped it were the builder's own read.
The engine now records `--tier`; this makes the prose choose.

- verify.md "Who refutes": T2 is the DEFAULT at floor ≥ plan — SPAWN
  `add-advisor` in refute mode, briefed from the frozen node before the
  diff, record its line with `--tier T2`; T1 is a prelude, never the
  rung's answer. The refute-read template now carries `--tier T2` (the
  T2 refuter of this task found the copy-paste path recorded an
  uncountable stamp). T0/T4 wording unchanged.
- SKILL.md step 3 VERIFY names the spawn before the gate, funded by
  trimming three cookbook comments (13256 B under the 13258 pin, 175
  lines, no budget literal moved; PROSE_PINS re-aimed).
- add-worker.md §4 verify bullet: spawn, record `--tier T2`, own read is
  T1. add-advisor.md's refute Return line carries `--tier T2`.
- docs/05 tier table: T2 "the default at floor ≥ plan — spawned, never
  optional"; T1 "a prelude".
- Refuted at T2 by a fresh add-advisor session: 3 probes held, 1 change
  acted on (recorded as `changed:`). The HEAD-vs-worktree checks in the
  new skill test declare themselves TRIPWIREs (R:SILENTTRIPWIRE).

Task: .add/tasks/t2-refute-default.md · milestone refute-at-t2
author: Tin Dang
Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
…ce mode — a notice, never a refusal

dogfood-and-measure F1: the session that wrote "at a plan floor a Must
carries two checks of different mode" froze 0 of 10 such Musts one task
later. Prose nobody follows binds nothing; a refusal would grow ceremony
back on the mechanical lane. So the freeze names them and judges nothing.

- `_check_lines` / `check_modes` / `single_mode_musts`: the mode word is
  the first token after the second `·` of a CHECKS line, from the closed
  set acceptance · property · contract · static · unit · e2e · manual,
  stripped of backticks/parentheses. Keyed by LINE, not by id.
- `freeze` at the refute rung's arming (standard|deep × floor plan|human
  × not explore) appends `notice: M1 (acceptance), M3 (contract) carry one
  evidence mode — a plan-floor Must carries two (direction.md § router)`
  to its success note; the stamp is written first. `add todo` carries the
  count. Quick · process floor · explore · Rejects · edges print nothing.
- FORMAT §6.2 states it; docs/03 and direction.md ×3 name it (line-neutral).
- Refuted at T2 three times by fresh add-advisor sessions: refute 1 FOUND
  a duplicated check id letting the last line's mode overwrite the first
  (fixed, graduated into edge E3, refrozen); refute 2 FOUND FORMAT §6.2
  documenting a notice shape the engine never emits (fixed, the prose
  check now matches the engine's format string); refute 3 recorded below.
- Direct fix (milestone exit 4): docs/03 no longer says acceptance mode
  is for "where no unit test fits" — every screen state is a filled edge
  with a bound acceptance check.

Engine twins ×4 synced; ENGINE_MD5 re-aimed (comment repaired to one
pointer). Pre-existing red at HEAD untouched: test_live_bundle_backfill_count.

Task: .add/tasks/two-mode-notice.md · milestone refute-at-t2
author: Tin Dang
Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
…o deltas folded

Every refute on this milestone was recorded at tier T2 by a spawned
add-advisor session: 6 refutes, 3 refuted (each graduated into an edge
and refrozen), 3 held; 5 of 18 probes changed the build or the spec.
The bench trigger does not fire. The single-mode notice fired on 3 of 3
freezes — the two-mode rule's first measured compliance is ~10%; the
promote-or-drop decision waits for n>3.

author: Tin Dang
Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
…nstead of reading it from the live bundle

test_live_bundle_backfill_count asserted the live bundle still carried
the pre-3.6 EVIDENCE/LESSONS scaffold, then synced and asserted it was
gone. c9b4feb backfilled the live bundle in the same commit, so the
precondition died and the tooling jobs on PR #223 went red on it.

The check now copies the bundle, re-seeds the scaffold into every done
task (117 on this bundle, floor 10), syncs, and asserts none remains.
Proven red by withholding the sync: 117 still carry a scaffold.

Quick lane · lesson filed under tdd.
author: Tin Dang
Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
… READMEs say what 3.6 ships

The tagline and the trust claims still described 3.5: "trust comes from
passing tests". Since evidence-over-tests and refute-at-t2 the gate
refuses a PASS nobody tried to refute at a plan-or-human floor, the
default frozen check for code is an acceptance check through a port,
property/contract/mutation checkers earn the same receipt, and the
freeze names a Must still on one kind of evidence.

- tagline: "Trust comes from evidence that survived a refute"
- 🔬 highlight: checks pass on a fresh receipt AND a session that did
  not build it tried to break the green; new 🎯 "Evidence, not test
  count" bullet, registered in the promised-capabilities guard
  (verb:refute · engine:CHECK_MODES)
- vanilla-vs-ADD "what you verify" row; How ADD Works Verify beat; the
  package README's flow line, non-negotiable 2, and the add-advisor
  refute-mode sentence

Identity lines untouched. Quick lane.
author: Tin Dang
Claude-Session: https://claude.ai/code/session_01NRFAPLzmF4GW5pGMCVXZUo
@TinDang97
TinDang97 merged commit 8999ae3 into main Sep 11, 2026
8 checks passed
@TinDang97
TinDang97 deleted the feat/evidence-over-tests branch September 11, 2026 02:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant