Skip to content

feat!: a judge that costs a seventh, measured on 228 real steps - #7

Merged
wasd96040501 merged 14 commits into
mainfrom
feat/cheaper-judge
Sep 23, 2026
Merged

wasd96040501 merged 14 commits into
mainfrom
feat/cheaper-judge

Conversation

@wasd96040501

Copy link
Copy Markdown
Owner

What this changes

  • The judge reads what says whether the work moves on, and little else (d6d746b): the requests (first and latest three whole), the assistant's latest messages, what its latest calls touched, its task list, and the step. No calls in full, no CLAUDE.md. About 2,200 tokens a judgement however long the session, down from ~17,000 and growing.
  • The judge no longer reads the step it judges twice (7fcf9c6). $.session.messages() already holds the step when it is judged, a block at a time; 0.8 left out only a last message matching the step whole, which never happens.
  • /taskcut and exact accounting of what the judge spends (4e7d40f). $.model.complete is not on the session's cost ledger, so /cost never showed the judge. taskcut now sums the API's own token counts in memory (free, no setting, nothing written), shows them in /taskcut, and ends each judgement's debug line with [judge sonnet: in=… cache_read=… cache_write=… out=… ms=…]. Claude Code's --debug is the debug mode; no taskcut-specific one.
  • floorPercent: 0 now judges from the first step (it waited for 5%).
  • Benchmark: make eval-replay (the judge over every step real sessions took, labelled, with intervals and prices; REF= for another commit's judge; replay-compare for paired differences), make eval-replay-check (the replay against a live session's prompts), replay-agree (label agreement), the click-zh workload (20 real click changes in one Chinese message, task list, a commit per issue), and the judge's cost in every end-to-end report.

Why

The judge's prompt was nine tenths commands in full and grew with the session: three cents a judgement, $2.57 over the stretch a 35% floor judges on a long run. And "is it no worse?" had no instrument that could answer it.

How it was tested

Claude Code 2.1.280, Sonnet 5. Full numbers and caveats: docs/measurement.md, "What the judge reads, replayed over real sessions".

The judge, 228 labelled steps from 7 committed sessions (82 boundaries, 129 not), 3 repeats, 95% intervals:

judge boundaries caught false NEXT $ / judgement
main (0.8), as it read live 99.2% (97.6–100) 5.2% (2.3–8.5) 0.0347
this PR 99.2% (97.6–100) 1.3% (0.3–2.8) 0.0050
this PR on haiku (negative control) 53.7% (45.5–61.0) 35.9% (29.2–42.6) 0.0020

Paired: boundaries +0.0% (−2.4 to +2.4), false NEXT −3.9% (−7.5 to −0.8). With one unpublished dev session added the false-NEXT interval is −6.2 to +0.2.

How the instrument itself was checked:

  • Replay fidelity: a recorder plugin beside taskcut in a real session; every step's judge prompt compared character for character with the replay's. First run 2/18 (found the double-read bug and the preserved-messages-after-compaction gap); after fixes 17/17, then 16/16 on a fresh session it was not tuned against.
  • Discrimination: haiku control separated by 45 points (recall) and 35 (false NEXT).
  • Labels: two blind independent annotators (Sonnet 5, guideline + sheets only). Cohen's kappa 0.87 / 0.89 vs committed, 0.91 between them; 21 splits settled by majority, 6 public labels changed, all recorded in eval/replay/disagreements.md. Please skim that file — the annotators share the judge's model family.
  • Harness tests (78): known-answer judges (oracle / always SAME / always NEXT / flip-flopping), check must accept/reject, TS↔Python contracts run under node, committed sets cover exactly the judged steps and hold no home directory.
  • Plugin: ./scripts/validate.sh (79 tests, tsc against regenerated 2.1.280 declarations, claude plugin validate).

End to end: sqlglot-long default — 36/36 issues, every check, 1 judgement (1,952 in / 47 out), 1 compaction. click-zh off/on — 23/23 checks each (never reached the floor). /taskcut checked in a real session against the debug log: identical.

  • ./scripts/validate.sh passes
  • tsc -p tsconfig.json passes against regenerated declarations (2.1.280)
  • Ran in a real interactive session
  • Claude Code version tested against: 2.1.280

Known gap this PR does not fix

A step that says nothing is never judged. In click-zh only 8 and 10 of 19 issue-to-issue moves had a judgeable step (the rest: a silent commit + the next issue's first command). And make eval-mechanism fails 1 of 9 checks in 2 of 2 runs today: the model did the second four-task turn without a word, so that turn got no compaction. Every judgement that was made was right. Likely fix for a follow-up: treat a step that updates the task list (task k completed, k+1 in progress) or commits as judgeable, and say so in the judge's rules — which needs new labels for those steps, so it is its own change.

Spend

About $49 at list price for this round's runs ($32 replay, of which $26 is the old judge; $9 sqlglot; $5 click-zh; the rest checks).

🤖 Generated with Claude Code

wasd96040501 and others added 14 commits September 23, 2026 14:24
The judge read what auto mode's permission classifier reads: every message
the person sent, every tool call in full except read-only lookups, and
CLAUDE.md. Replayed over real sessions, nine tenths of that prompt was the
calls in full, and they are what made it grow with the session -- sixteen
thousand tokens a judgement on average, three cents -- while saying least
about whether a piece of work is done.

It now reads the step; every request, the first and the latest three whole
and the rest cut to a line; the assistant's latest ten messages; what its
latest twelve calls touched, a file or what a command says it does, never
the call in full; and its task list. Never any tool output, and not
CLAUDE.md, so taskcut no longer reads any file. The prompt is about two
thousand tokens however long the session has run.

Messages nobody typed -- a background task's report, a plugin's prompt,
taskcut's own Continue. -- are not read as requests, and a compaction's
summary is read as one.

BREAKING CHANGE: the judge no longer reads CLAUDE.md or tool calls in full,
and judgePrompt takes no memory argument.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The bench hook takes its prompt and drops it. When the hook fails instead --
it throws, or the engine refuses it -- the prompt goes on to the model as an
ordinary request, and with tools a model sets about working out what
"taskcut-judgebench <paths>" means, in the scratch directory, unattended. It
now runs with --tools "".

An unanswered call recorded only "api-error", so a spent weekly limit read the
same as a refused request. The bench now records the engine's status and
error beside it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A plugin's model calls are nobody's to count but the plugin's: on Claude Code
2.1.280, $.model.complete is not on the session's cost ledger -- a call made
from a hook moves $.session.usage().cost by nothing, so /cost and the status
line leave the judge out -- and the transcript does not record it either.

Each call does resolve the API's own token counts, so taskcut now sums them
in memory, for free, and shows them two ways, with no setting:

* /taskcut, an immediate command, gives the session's totals: steps judged,
  tokens read and written, compactions started and their sizes.
* Each judgement's debug line ends with a record of the call,
  [judge sonnet: in=1834 cache_read=0 cache_write=0 out=52 ms=2140], which
  the benchmark adds up. Claude Code's --debug is the debug mode; taskcut
  adds none of its own.

Tokens, not dollars: a price is the host's to know. Hooking /cost itself was
tried: interactively it opens the usage panel, which prints no text a
command.run hook can extend.

An API error is now logged with its status and kind, and a compaction the
engine skipped is no longer counted as one.

Fixed on the way, found by the check of /taskcut: a floorPercent of 0 judged
nothing until the context reached 5%. The wait after a compaction that could
not get under the floor read "no compaction yet" as one that left 0%. The
rule is now judgingFrom in config.ts, with tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The benchmark read a run's cost off its transcript, and the judge's calls are
not in it -- nor in Claude Code's cost ledger -- so every cost it has
reported for a taskcut arm left the judging out.

Every session now runs with --debug-file. After the run, the record taskcut
puts at the end of each judgement line is summed into
results/<run>.judging.json, beside the transcript, and the cost table gains
the judge's calls, weighted input and output. The debug log itself stays in
the scratch directory.

The format is a contract between hooks/spend.ts and metrics.JUDGEMENT_LINE,
so one test renders a line with the plugin's own code, under node, and
parses it here.

Checked on a real `make eval-mechanism` run: nine of nine checks pass, and
the sidecar holds its three judgements, 5,933 tokens in and 161 out.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every boundary the judge has been measured on came from sqlglot-long: one
repository, one prompt, one way of saying "Issue 12 done; now 13", in
sessions that never kept a task list and never committed. Real sessions do
both, and a commit is the most common thing a step does between one piece of
work and the next -- wrapping up, not moving on.

click-zh hands the twenty click changes of issues-long over in one message,
in Chinese: keep a task list, work through ISSUES.md in order, commit after
each issue. It is graded by the same per-issue tests and the whole suite.

A session that commits moves HEAD, so `git diff HEAD -- tests` no longer
shows whether a test was touched. click_issues.py now tags the broken state
eval-start, writes ISSUES.md for a workload that names itself, and the
checks diff against the tag; one more asks that the work was committed.

Verified both ways: in the broken state every issue check fails; with the
twenty real fixes committed one by one every check passes; a test edited
and committed fails tests_untouched.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On Claude Code 2.1.280 TaskCreate and TaskUpdate are off in a session unless
CLAUDE_CODE_ENABLE_TODO_TOOLS is set, and the first click-zh run spent its
opening steps searching for them. A workload can now set environment
variables for its sessions, the same under every arm, and click-zh sets that
one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a step is judged, $.session.messages() (2.1.280) already holds it, and
not as one message: a response is held a block at a time -- its words, then
each call -- followed by the result of any call that has already run.
beforeStep left out only a last message matching the step whole, which
never happens: the step was read twice, once as the latest thing the
assistant said, and whether it was depended on how fast its tools ran.

The step is now the shortest run of messages at the end whose words, joined,
are the step's and whose calls are the step's, with nothing between them but
those calls' results; it is left out, and nothing else is.

Found by the replay check: a live session with a recorder beside taskcut,
the judge's prompt at every step compared with the one the replay rebuilds
from the transcript. 2 of 18 steps matched before, 17 of 17 after (with the
replay holding messages as the engine does; next commit).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
make eval-judge asks the judge about sixteen steps written for it. make
eval-replay asks it about every step taskcut would have judged in real
sessions, as the session looked then, three times each, and reports
boundaries caught and false NEXT with 95% bootstrap intervals, hard SAME,
agreement between repeats, and list price per judgement and for the
stretch of a long session past the floor. REF= measures hooks/judge.ts as
it was at a commit, and replay-compare pairs two runs step by step, with an
interval on the difference -- what "not worse" has to rest on.

The sets are committed in eval/replay: the sqlglot-long, issues-long and
mechanism sessions the labels were made on, and the two new click-zh ones,
scrubbed of the home directory. Every step taskcut judges is labelled N, S
or E (LABELS.md); a set with a step unlabelled is refused. Sessions that are
someone's own go in eval/replay/local, which is not committed.

The replay is checked against what the judge is given live:
make eval-replay-check runs a real session with a recorder plugin beside
taskcut and compares, character for character, the prompt taskcut builds
at every step with the one the replay rebuilds. Its first run matched 2 of
18 steps. The engine holds a response a block at a time, and after a
compaction keeps the latest messages whole behind the summary (named in
the boundary's preservedMessages); the replay now does both, and matches
17 of 17 steps across two compactions, a task list and the Continue. prompt.

The harness is tested against judges whose answers are known -- always
right, always SAME, always NEXT, one that changes its mind -- against
records the check must accept and must reject, and, where Python mirrors
the plugin, against the plugin's own TypeScript run under node.

The bench now reads a case's messages out of a set by position (a set holds
a session once; $.fs.read takes nothing over 4 MiB), retries rate limits
and overloads, and records each call's latency.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re label agreement

A bench case now ends where the step's own messages end, before its calls'
results, as $.session.messages() holds it when the step is judged: a judge
has to leave the step out itself, as it has to live. The judge up to 0.8
did not, and the replay now measures what it was actually given.

replay-agree compares the committed labels with independent labellings of
the same sets: Cohen's kappa pairwise and per set, and every step on which
they do not all agree, with each labeller's reason.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
make eval-replay, three repeats each, Sonnet 5, Claude Code 2.1.280:

                    boundaries caught   false NEXT          $ a judgement
  main              99.2% (96.8-100)    4.3% (2.0-7.2)      0.0342
  this branch       99.2% (97.6-100)    1.1% (0.2-2.5)      0.0052

Paired over the same 83 boundaries and 149 steps that are not one: +0.0%
(-2.4 to +2.4) on boundaries, -3.1% (-6.0 to -0.4) on false NEXT, at 0.15x
the price. main's judge is given the step as the engine holds it, which it
read twice.

Each run now works in its own scratch directory, so two judges can be
measured at once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The same judge on Haiku 4.5, three repeats over the 257 steps: 54.2% of
boundaries caught (46.2-61.4) and 33.8% false NEXT (27.7-40.0), against
99.2% and 1.1% on Sonnet 5. A benchmark that could not tell these apart
could not be trusted with the smaller differences it is asked about.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every set was labelled again, blind, by two annotators (Claude Sonnet 5)
given only LABELS.md and the sheets. Cohen's kappa over the 257 steps:
0.87 and 0.89 against the committed labels, 0.91 between the two. On the
21 steps they did not all agree, two of three decide, and E where all
three differ; six public labels change -- sqlglot-long--off 81 and 82 had
the checkpoint and the move the wrong way round, against LABELS.md's own
example -- and eval/replay/disagreements.md records every split with the
annotators' reasons.

Rescored on the settled labels, public sets (the committed runs, no new
calls): boundaries caught 99.2% for both judges; false NEXT 5.2% for
main's and 1.3% for this one, a difference of -3.9% (-7.5 to -0.8).

Fixed on the way: a scored subset counted and priced every set's calls,
and make eval-report failed on a run with a judging sidecar (Run is
frozen); a test now goes through the report command.

sqlglot-long, arm default, end to end: all 36 issues and every check pass;
taskcut judged one step, 1,952 tokens in and 47 out, and compacted once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
runs.json was rewritten from the transcripts in the results directory, and
a checkout rarely has them all -- they are not committed -- so reporting a
new run dropped every other run's numbers. The runs whose transcripts are
there now update their entries, and every other entry is kept.

With it, the click-zh arms and the new sqlglot-long default run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
measurement.md gets the replay: this change against 0.8 on 228 real steps,
with intervals, the Haiku control, the check that the replay is what the
judge sees, how far three labellers agreed, and what it all cost.

It also says what the benchmark does not cover, and one thing it found that
this change does not fix: a step that says nothing is never judged. In
click-zh only 8 and 10 of 19 moves between issues came with a step taskcut
would judge, and make eval-mechanism now fails one check twice out of two,
on a turn of four tasks the model did without a word. The README's
limitations and troubleshooting say so.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@wasd96040501
wasd96040501 merged commit 4bf4389 into main Sep 23, 2026
3 checks passed
@wasd96040501
wasd96040501 deleted the feat/cheaper-judge branch September 23, 2026 16:15
@wasd96040501 wasd96040501 mentioned this pull request Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant