Skip to content

docs(release): name the per-docket-Term skill anchor in §3; add segment-anchors - #2229

Merged
modelmirror merged 3 commits into
stagingfrom
fix/release-skill-anchor
Oct 3, 2026
Merged

modelmirror merged 3 commits into
stagingfrom
fix/release-skill-anchor

Conversation

@modelmirror

@modelmirror modelmirror commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #2226 (Release 1, #1993). A docs fix, a read-only command, and the maintainer-decided correction of the registered surfaces.

Findings

Root cause. The pipeline scores correctly. The wrong part is the quoted figure. A cert cell's skill anchor is the risk-set band rate, pooled over the statpack Terms strictly before the prediction's frozen context.term. That Term is the docket-number Term. The definition is the same in each of these places:

  • pipeline/base_rates.py _pooled_band_rate / prediction_base_rate
  • cell_context setting term = scotus_term_year(docket_number)
  • the evaluate and predict prompts
  • the statpack footnote

The cohort decided at the OT2026 long conference is almost all 25- dockets: about 108 of the 110 registered cert/distribution events, and all 10 CVSG events. So its anchor is the OT2017–OT2024 pool. The 5.02 / 16.89 / 35.51 / 70.79 / 23.63 figures quoted in docs/freeze-record.md, metrics/README.md and §3 are the OT2026-docket pool (OT2017–OT2025). They are the registered rule applied to the wrong docket Term.

Numbers reproduced. Committed metrics/statpack.json, identical on origin/main (808f812) and the three prior builds; sal-v4; lookback 10. The lookback does not bind because the pack starts at OT2017.

docket Term pools baseline elevated high federal state
2025 OT2017–OT2024 (8 Terms) 5.121% 17.224% (484/2810) 34.967% 72.928% 22.704%
2026 OT2017–OT2025 (9 Terms) 5.016% 16.888% 35.507% 70.792% 23.628%

How each recorded grading value arises. These are the 9 proc-v8 risk_set gradings on main, all on scotus/73281619, which is elevated with context.term=2025:

  • claude-judge 0.172242: the exact pool.
  • gemini-judge 0.1722: the same pool at four decimals.
  • codex-judge 0.172379359430605: exactly the n-weighted pool of the statpack.md rounded reached rows for 2017–2024. The evaluate prompt points evaluators at that table, so a spread of up to about 0.0005 is by design.

The issue's figure of about 16.90% is the same rounded-row pool taken over 2017–2025.

Quoted registrations.

  • docs/freeze-record.md ~L3832: "the risk-set (reached) rate pooled over base_rate_lookback_terms, excluding the cell's own Term … On the committed statpack that is baseline 5.02% / elevated 16.89% / high 35.51% / federal 70.79% / state 23.63%"
  • ~L4952: "the registered band rates are the skill anchor only — 5.02% baseline / 16.89% elevated / …"
  • metrics/README.md L74 and L566 quote the same set.

The definition registered there ("excluding the cell's own Term") is the one the code implements. The numbers quoted beside it are that definition evaluated for an OT2026 docket.

All bands, not just elevated. Moving from the OT2026-docket pool to the OT2025-docket pool changes each band as follows: baseline +0.10 pp, elevated +0.34 pp, high −0.54 pp, federal +2.14 pp, state −0.92 pp. Band-mix implied grant rate at the OT2025 rates:

  • 10.24% over 110 events (registered: 10.05%)
  • 12.30% over 120 events (registered: 12.17%)
  • selected subset 18.18% (n=39; registered: 17.83%)
  • declined 5.88% (n=71; registered: 5.78%)

Consequence: no re-grade owed. The leaderboard reads each grading's own recorded segment_base_rate, both for the prior-Term skill (leaderboard.py _prior_baseline) and for grants_expected. No code, prompt or config quotes 16.89%. stats-reviewer independently checked all 9 frozen risk_set cert gradings: each sits within 6e-4 of the docket-Term pool for its scored prediction's (term, band). Leakage: pooling OT2025 for a 25- docket would be wrong twice over. That row contains the case itself, and it is right-censored: elevated reached is 13.5% on n=275, against 13.8–20.5% in the mature Terms.

What changed

  • New read-only command fedcourts segment-anchors --term N [--term M] in src/fedcourtsai/cli.py. It reads the committed metrics/statpack.json and the salience config, never the corpus. For each band it prints the risk_set and terminal rates through the scorer's own _pooled_band_rate, the weighted resolved n, and the Terms that contributed. Output is JSON on stdout and a table on stderr. Tests are in tests/test_cli_segment_anchors.py, and there is a docs/cli.md row. No scoring code changed.
  • docs/release-ot2026-long-conference.md §3 only:
    • The anchor is defined by docket Term.
    • The Evidence block runs uv run fedcourts segment-anchors --term 2025 --term 2026.
    • The Prose block quotes both docket-Term pools and says that skill is computed against each grading's recorded rate, with the rounding spread.
    • The band-mix figures are corrected to 10.2% / 18.2% (n=39) / 5.9% (n=71) / 12.3%.
    • Each band row must count its scored events per docket Term at release time. Post-freeze additions are mostly OT2026 dockets.
    • §2/§5 are untouched; another agent owns them.

Produced output on the committed pack:

docket Term 2025 (sal-v4, lookback 10; pools OT2017-OT2024, 8 Term(s)):
  federal   risk_set 72.93% (n=181)  terminal 72.93% (n=181)
  high      risk_set 34.97% (n=898)  terminal 34.97% (n=898)
  state     risk_set 22.70% (n=392)  terminal 16.91% (n=278)
  elevated  risk_set 17.22% (n=2810)  terminal 10.46% (n=2026)
  baseline  risk_set 5.12% (n=11580)  terminal 1.24% (n=8770)
docket Term 2026 (sal-v4, lookback 10; pools OT2017-OT2025, 9 Term(s)):
  federal   risk_set 70.79% (n=202)  ...
  elevated  risk_set 16.89% (n=3085)  ...

Maintainer decision (2026-10-03): registered surfaces corrected in this PR

No re-grade and no registered-rate change: the rule is unchanged, only the quoted figures move.

  • docs/freeze-record.md: a correction entry is appended at the end. Landed entries are untouched. The entry:
    • cites the 2026-09-15 entry's anchor bullet and the 2026-09-24 entry's "registered band rates are the skill anchor only" bullet;
    • states that the definition is unchanged and is what the code implements;
    • gives both docket-Term pools and the corrected band-mix figures (10.24 / 12.30 / 18.18 at n=39 / 5.88 at n=71, replacing 10.05 (~10.1) / ~12.2 / ~17.8 / ~5.8);
    • notes that the 09-15 entry's 8.30% / 6.67% comparison rates share the old basis (the in-scope 180 re-reads at 8.44%, still about 1.2x; the distributed set is re-read at release);
    • records that nothing moved (all 9 frozen gradings within 6e-4; digests unchanged);
    • carries a segment-anchors effect check and <FILL: …> promotion placeholders.
  • metrics/README.md: both quotes now state the per-docket-Term definition within the lookback, quote both pools (the OT2026 one noted as moving with each build), and point to segment-anchors.
  • §3 is aligned with both.

Reviews and gate

  • stats-reviewer: no blockers.

    • Applied: R3 (count the docket-Term mix per row, disclose that the OT2026 anchor moves with the build); R4 (exact subset figures with their n); N1 (list only the Terms that contributed); N2 (the command prints the exact pool, the recorded rate can differ by ≤0.0005); N3 (name which value is the rounded pool and which the exact).
    • R1/R2 (freeze record, metrics/README) are deferred as above.
  • code-reviewer: no blockers.

    • Applied: reuse the scorer's _version_segments rather than a copy; tests for rate-less slices, Terms without the pinned version, and equality with a direct _pooled_band_rate call; docstring wording.
    • Not applied: refactoring the pooler into a public helper. That would change scoring code.
  • docs-reviewer: no blockers.

    • Applied: the judge-rounding sentence matched to the committed gradings; "registered" limited to the rule (the OT2026 figures are called "the freeze-record figures"); cli.md names the config read.
    • metrics/README is deferred as above.
  • Second round, on the correction commits: stats-reviewer and docs-reviewer found no blockers. Applied:

    • quote the replaced figures as printed;
    • address the 09-15 entry's comparison rates;
    • say "the OT2026-docket pool" in §3;
    • add the lookback bound and the moving-with-build caveat to metrics/README;
    • reflow.

    Not applied (optional): restating the complements, which no reading uses as floors. The entry says they move with the rates.

  • scripts/gate.sh lint types schemas test: all green. 6075 passed, 2 skipped. The full suite ran before the final reviewer edits. Those edits touched only the new command and its test file. Lint, types and the command's own tests were re-run on the final tree.

Leaving this for the maintainer to merge (release figures plus a freeze-record entry). Rebased on current staging, after #2230; §1/§2/§5 untouched. The affected tests (test_process_version, test_arrival_membership_record, test_statpack, test_evaluate, test_cli_segment_anchors, test_workflow_auth_gate) passed after the rebase, and lint is clean.

🤖 Generated with Claude Code

modelmirror and others added 3 commits October 3, 2026 22:54
…nt-anchors

The cert skill anchor pools statpack Terms strictly before the prediction's
frozen docket-number Term. The cohort is almost entirely OT2025 dockets, so
it is scored against the OT2017-OT2024 pool (elevated 17.22%), not the
OT2026-docket pool (16.89%) that §3 quoted. Add a read-only
`fedcourts segment-anchors` command that prints the pooled per-band anchors
per docket Term through the scorer's own pooler, and point §3 at it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…entry, metrics/README)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s' basis

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@modelmirror
modelmirror force-pushed the fix/release-skill-anchor branch from 3ac279e to d68497f Compare October 3, 2026 23:00
@modelmirror
modelmirror merged commit 460790e into staging Oct 3, 2026
9 of 12 checks passed
@modelmirror
modelmirror deleted the fix/release-skill-anchor branch October 3, 2026 23:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant