docs(release): name the per-docket-Term skill anchor in §3; add segment-anchors - #2229
Merged
Merged
Conversation
…nt-anchors The cert skill anchor pools statpack Terms strictly before the prediction's frozen docket-number Term. The cohort is almost entirely OT2025 dockets, so it is scored against the OT2017-OT2024 pool (elevated 17.22%), not the OT2026-docket pool (16.89%) that §3 quoted. Add a read-only `fedcourts segment-anchors` command that prints the pooled per-band anchors per docket Term through the scorer's own pooler, and point §3 at it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…entry, metrics/README) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s' basis Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
modelmirror
force-pushed
the
fix/release-skill-anchor
branch
from
October 3, 2026 23:00
3ac279e to
d68497f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #2226 (Release 1, #1993). A docs fix, a read-only command, and the maintainer-decided correction of the registered surfaces.
Findings
Root cause. The pipeline scores correctly. The wrong part is the quoted figure. A cert cell's skill anchor is the risk-set band rate, pooled over the statpack Terms strictly before the prediction's frozen
context.term. That Term is the docket-number Term. The definition is the same in each of these places:pipeline/base_rates.py_pooled_band_rate/prediction_base_ratecell_contextsettingterm = scotus_term_year(docket_number)The cohort decided at the OT2026 long conference is almost all
25-dockets: about 108 of the 110 registered cert/distribution events, and all 10 CVSG events. So its anchor is the OT2017–OT2024 pool. The 5.02 / 16.89 / 35.51 / 70.79 / 23.63 figures quoted indocs/freeze-record.md,metrics/README.mdand §3 are the OT2026-docket pool (OT2017–OT2025). They are the registered rule applied to the wrong docket Term.Numbers reproduced. Committed
metrics/statpack.json, identical on origin/main (808f812) and the three prior builds; sal-v4; lookback 10. The lookback does not bind because the pack starts at OT2017.How each recorded grading value arises. These are the 9 proc-v8
risk_setgradings on main, all onscotus/73281619, which is elevated withcontext.term=2025:reachedrows for 2017–2024. The evaluate prompt points evaluators at that table, so a spread of up to about 0.0005 is by design.The issue's figure of about 16.90% is the same rounded-row pool taken over 2017–2025.
Quoted registrations.
docs/freeze-record.md~L3832: "the risk-set (reached) rate pooled overbase_rate_lookback_terms, excluding the cell's own Term … On the committed statpack that is baseline 5.02% / elevated 16.89% / high 35.51% / federal 70.79% / state 23.63%"metrics/README.mdL74 and L566 quote the same set.The definition registered there ("excluding the cell's own Term") is the one the code implements. The numbers quoted beside it are that definition evaluated for an OT2026 docket.
All bands, not just elevated. Moving from the OT2026-docket pool to the OT2025-docket pool changes each band as follows: baseline +0.10 pp, elevated +0.34 pp, high −0.54 pp, federal +2.14 pp, state −0.92 pp. Band-mix implied grant rate at the OT2025 rates:
Consequence: no re-grade owed. The leaderboard reads each grading's own recorded
segment_base_rate, both for the prior-Term skill (leaderboard.py_prior_baseline) and forgrants_expected. No code, prompt or config quotes 16.89%. stats-reviewer independently checked all 9 frozenrisk_setcert gradings: each sits within 6e-4 of the docket-Term pool for its scored prediction's(term, band). Leakage: pooling OT2025 for a25-docket would be wrong twice over. That row contains the case itself, and it is right-censored: elevatedreachedis 13.5% on n=275, against 13.8–20.5% in the mature Terms.What changed
fedcourts segment-anchors --term N [--term M]insrc/fedcourtsai/cli.py. It reads the committedmetrics/statpack.jsonand the salience config, never the corpus. For each band it prints therisk_setandterminalrates through the scorer's own_pooled_band_rate, the weighted resolvedn, and the Terms that contributed. Output is JSON on stdout and a table on stderr. Tests are intests/test_cli_segment_anchors.py, and there is adocs/cli.mdrow. No scoring code changed.docs/release-ot2026-long-conference.md§3 only:uv run fedcourts segment-anchors --term 2025 --term 2026.Produced output on the committed pack:
Maintainer decision (2026-10-03): registered surfaces corrected in this PR
No re-grade and no registered-rate change: the rule is unchanged, only the quoted figures move.
docs/freeze-record.md: a correction entry is appended at the end. Landed entries are untouched. The entry:segment-anchorseffect check and<FILL: …>promotion placeholders.metrics/README.md: both quotes now state the per-docket-Term definition within the lookback, quote both pools (the OT2026 one noted as moving with each build), and point tosegment-anchors.Reviews and gate
stats-reviewer: no blockers.
code-reviewer: no blockers.
_version_segmentsrather than a copy; tests for rate-less slices, Terms without the pinned version, and equality with a direct_pooled_band_ratecall; docstring wording.docs-reviewer: no blockers.
Second round, on the correction commits: stats-reviewer and docs-reviewer found no blockers. Applied:
Not applied (optional): restating the complements, which no reading uses as floors. The entry says they move with the rates.
scripts/gate.sh lint types schemas test: all green. 6075 passed, 2 skipped. The full suite ran before the final reviewer edits. Those edits touched only the new command and its test file. Lint, types and the command's own tests were re-run on the final tree.Leaving this for the maintainer to merge (release figures plus a freeze-record entry). Rebased on current
staging, after #2230; §1/§2/§5 untouched. The affected tests (test_process_version,test_arrival_membership_record,test_statpack,test_evaluate,test_cli_segment_anchors,test_workflow_auth_gate) passed after the rebase, and lint is clean.🤖 Generated with Claude Code