Skip to content

155 of 158 scripts/** self-tests have no assertion floor: a battery that never ran is indistinguishable from one that passed #13799

Description

@claude

Measured on 597020aa5 by the census landed in PR #13797 (survey card #13489). Filed unassigned.

The reading

155 of 158 scripts/** self-tests decide success by failures.length === 0 and nothing else, so "every case held" and "the cases never ran" print the same line. One had a floor at the time of the survey (scripts/check-doc-authoring.mjs, from PR #13487); PR #13797 adds two more.

⚠️ A verdict line that already prints a case count is evidence, not proof. Two files here derive and print a SELF_TEST_CASE_COUNT that nothing ever compares — scripts/typecheck-configs.mjs and scripts/check-comment-mask-corpus.mjs. If a case array shrinks, the printed number shrinks with it and the gate stays green. ⛔ Pinning a TOTAL is not the repair either: it rots the moment a sibling battery grows.

This is hole 1 of #13489, ⛔ orthogonal to hole 2 (no verdict handshake, tracked separately). ⛔ Do not merge the two counts.

The remedy

The PR #13487 shape, pinning registered names rather than numbers: a frozen roster of battery names each with its own case floor, every battery opening with battery('...'), every assertion attributed to the battery most recently opened, a refusal when the opened set differs from the declared set, and the roster's own size pinned too — deleting an entry silences a floor exactly as effectively as zeroing it. A set difference says which battery stopped; a count says only that something did.

Tiers, by what the transplant actually costs

  • Tier B — direct. Self-tests already sectioned into named groups (most carry section banner comments, or block-scoped groups). The roster transplants verbatim; no case is rewritten. Both gates fixed in PR Survey: which scripts/** self-tests cannot prove they ran — and the two that now can #13797 were Tier B.
  • Tier C — reshaping first. 3 inline top-level self-test blocks with no entry function (scripts/check-regen-pending.mjs, scripts/git-merge-regen.mjs, scripts/setup-git-hooks.mjs); 2 multi-entry dispatches combining several self-test callees (scripts/check-platform-checklist.mjs, scripts/check-durability-degradation-log-level.mjs); and table-driven self-tests where the natural roster is the table's own named rows rather than sections.

One caveat on the number, from the instrument's own header

The static criterion reads names. A floor spelled with names it does not know reads as unfloored, so its error runs in one direction and 155 is an upper bound. The tree was independently hand-swept for zero-case refusals outside those names; every hit was a production-scan refusal, not a self-test floor. Re-measure with node scripts/measure-self-test-floor.mjs rather than re-deriving.

A floor is not always the measured count

Recorded in PR #13797 and worth carrying forward: where a battery's case count is one-per-row of a ⛔ shrink-only ledger, a floor at today's count reddens every legitimate shrink and trains the next author to edit the floor — the one habit these floors exist to prevent. Pin the part that does not move with the list instead (the structural case ran, and at least one row was audited), and say so in place.


Generated by Claude Code

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions