Repository navigation
feat!: a judge that costs a seventh, measured on 228 real steps - #7
Merged
Merged
Conversation
The judge read what auto mode's permission classifier reads: every message the person sent, every tool call in full except read-only lookups, and CLAUDE.md. Replayed over real sessions, nine tenths of that prompt was the calls in full, and they are what made it grow with the session -- sixteen thousand tokens a judgement on average, three cents -- while saying least about whether a piece of work is done. It now reads the step; every request, the first and the latest three whole and the rest cut to a line; the assistant's latest ten messages; what its latest twelve calls touched, a file or what a command says it does, never the call in full; and its task list. Never any tool output, and not CLAUDE.md, so taskcut no longer reads any file. The prompt is about two thousand tokens however long the session has run. Messages nobody typed -- a background task's report, a plugin's prompt, taskcut's own Continue. -- are not read as requests, and a compaction's summary is read as one. BREAKING CHANGE: the judge no longer reads CLAUDE.md or tool calls in full, and judgePrompt takes no memory argument. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The bench hook takes its prompt and drops it. When the hook fails instead -- it throws, or the engine refuses it -- the prompt goes on to the model as an ordinary request, and with tools a model sets about working out what "taskcut-judgebench <paths>" means, in the scratch directory, unattended. It now runs with --tools "". An unanswered call recorded only "api-error", so a spent weekly limit read the same as a refused request. The bench now records the engine's status and error beside it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A plugin's model calls are nobody's to count but the plugin's: on Claude Code 2.1.280, $.model.complete is not on the session's cost ledger -- a call made from a hook moves $.session.usage().cost by nothing, so /cost and the status line leave the judge out -- and the transcript does not record it either. Each call does resolve the API's own token counts, so taskcut now sums them in memory, for free, and shows them two ways, with no setting: * /taskcut, an immediate command, gives the session's totals: steps judged, tokens read and written, compactions started and their sizes. * Each judgement's debug line ends with a record of the call, [judge sonnet: in=1834 cache_read=0 cache_write=0 out=52 ms=2140], which the benchmark adds up. Claude Code's --debug is the debug mode; taskcut adds none of its own. Tokens, not dollars: a price is the host's to know. Hooking /cost itself was tried: interactively it opens the usage panel, which prints no text a command.run hook can extend. An API error is now logged with its status and kind, and a compaction the engine skipped is no longer counted as one. Fixed on the way, found by the check of /taskcut: a floorPercent of 0 judged nothing until the context reached 5%. The wait after a compaction that could not get under the floor read "no compaction yet" as one that left 0%. The rule is now judgingFrom in config.ts, with tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The benchmark read a run's cost off its transcript, and the judge's calls are not in it -- nor in Claude Code's cost ledger -- so every cost it has reported for a taskcut arm left the judging out. Every session now runs with --debug-file. After the run, the record taskcut puts at the end of each judgement line is summed into results/<run>.judging.json, beside the transcript, and the cost table gains the judge's calls, weighted input and output. The debug log itself stays in the scratch directory. The format is a contract between hooks/spend.ts and metrics.JUDGEMENT_LINE, so one test renders a line with the plugin's own code, under node, and parses it here. Checked on a real `make eval-mechanism` run: nine of nine checks pass, and the sidecar holds its three judgements, 5,933 tokens in and 161 out. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every boundary the judge has been measured on came from sqlglot-long: one repository, one prompt, one way of saying "Issue 12 done; now 13", in sessions that never kept a task list and never committed. Real sessions do both, and a commit is the most common thing a step does between one piece of work and the next -- wrapping up, not moving on. click-zh hands the twenty click changes of issues-long over in one message, in Chinese: keep a task list, work through ISSUES.md in order, commit after each issue. It is graded by the same per-issue tests and the whole suite. A session that commits moves HEAD, so `git diff HEAD -- tests` no longer shows whether a test was touched. click_issues.py now tags the broken state eval-start, writes ISSUES.md for a workload that names itself, and the checks diff against the tag; one more asks that the work was committed. Verified both ways: in the broken state every issue check fails; with the twenty real fixes committed one by one every check passes; a test edited and committed fails tests_untouched. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On Claude Code 2.1.280 TaskCreate and TaskUpdate are off in a session unless CLAUDE_CODE_ENABLE_TODO_TOOLS is set, and the first click-zh run spent its opening steps searching for them. A workload can now set environment variables for its sessions, the same under every arm, and click-zh sets that one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a step is judged, $.session.messages() (2.1.280) already holds it, and not as one message: a response is held a block at a time -- its words, then each call -- followed by the result of any call that has already run. beforeStep left out only a last message matching the step whole, which never happens: the step was read twice, once as the latest thing the assistant said, and whether it was depended on how fast its tools ran. The step is now the shortest run of messages at the end whose words, joined, are the step's and whose calls are the step's, with nothing between them but those calls' results; it is left out, and nothing else is. Found by the replay check: a live session with a recorder beside taskcut, the judge's prompt at every step compared with the one the replay rebuilds from the transcript. 2 of 18 steps matched before, 17 of 17 after (with the replay holding messages as the engine does; next commit). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
make eval-judge asks the judge about sixteen steps written for it. make eval-replay asks it about every step taskcut would have judged in real sessions, as the session looked then, three times each, and reports boundaries caught and false NEXT with 95% bootstrap intervals, hard SAME, agreement between repeats, and list price per judgement and for the stretch of a long session past the floor. REF= measures hooks/judge.ts as it was at a commit, and replay-compare pairs two runs step by step, with an interval on the difference -- what "not worse" has to rest on. The sets are committed in eval/replay: the sqlglot-long, issues-long and mechanism sessions the labels were made on, and the two new click-zh ones, scrubbed of the home directory. Every step taskcut judges is labelled N, S or E (LABELS.md); a set with a step unlabelled is refused. Sessions that are someone's own go in eval/replay/local, which is not committed. The replay is checked against what the judge is given live: make eval-replay-check runs a real session with a recorder plugin beside taskcut and compares, character for character, the prompt taskcut builds at every step with the one the replay rebuilds. Its first run matched 2 of 18 steps. The engine holds a response a block at a time, and after a compaction keeps the latest messages whole behind the summary (named in the boundary's preservedMessages); the replay now does both, and matches 17 of 17 steps across two compactions, a task list and the Continue. prompt. The harness is tested against judges whose answers are known -- always right, always SAME, always NEXT, one that changes its mind -- against records the check must accept and must reject, and, where Python mirrors the plugin, against the plugin's own TypeScript run under node. The bench now reads a case's messages out of a set by position (a set holds a session once; $.fs.read takes nothing over 4 MiB), retries rate limits and overloads, and records each call's latency. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re label agreement A bench case now ends where the step's own messages end, before its calls' results, as $.session.messages() holds it when the step is judged: a judge has to leave the step out itself, as it has to live. The judge up to 0.8 did not, and the replay now measures what it was actually given. replay-agree compares the committed labels with independent labellings of the same sets: Cohen's kappa pairwise and per set, and every step on which they do not all agree, with each labeller's reason. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
make eval-replay, three repeats each, Sonnet 5, Claude Code 2.1.280:
boundaries caught false NEXT $ a judgement
main 99.2% (96.8-100) 4.3% (2.0-7.2) 0.0342
this branch 99.2% (97.6-100) 1.1% (0.2-2.5) 0.0052
Paired over the same 83 boundaries and 149 steps that are not one: +0.0%
(-2.4 to +2.4) on boundaries, -3.1% (-6.0 to -0.4) on false NEXT, at 0.15x
the price. main's judge is given the step as the engine holds it, which it
read twice.
Each run now works in its own scratch directory, so two judges can be
measured at once.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The same judge on Haiku 4.5, three repeats over the 257 steps: 54.2% of boundaries caught (46.2-61.4) and 33.8% false NEXT (27.7-40.0), against 99.2% and 1.1% on Sonnet 5. A benchmark that could not tell these apart could not be trusted with the smaller differences it is asked about. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every set was labelled again, blind, by two annotators (Claude Sonnet 5) given only LABELS.md and the sheets. Cohen's kappa over the 257 steps: 0.87 and 0.89 against the committed labels, 0.91 between the two. On the 21 steps they did not all agree, two of three decide, and E where all three differ; six public labels change -- sqlglot-long--off 81 and 82 had the checkpoint and the move the wrong way round, against LABELS.md's own example -- and eval/replay/disagreements.md records every split with the annotators' reasons. Rescored on the settled labels, public sets (the committed runs, no new calls): boundaries caught 99.2% for both judges; false NEXT 5.2% for main's and 1.3% for this one, a difference of -3.9% (-7.5 to -0.8). Fixed on the way: a scored subset counted and priced every set's calls, and make eval-report failed on a run with a judging sidecar (Run is frozen); a test now goes through the report command. sqlglot-long, arm default, end to end: all 36 issues and every check pass; taskcut judged one step, 1,952 tokens in and 47 out, and compacted once. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
runs.json was rewritten from the transcripts in the results directory, and a checkout rarely has them all -- they are not committed -- so reporting a new run dropped every other run's numbers. The runs whose transcripts are there now update their entries, and every other entry is kept. With it, the click-zh arms and the new sqlglot-long default run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
measurement.md gets the replay: this change against 0.8 on 228 real steps, with intervals, the Haiku control, the check that the replay is what the judge sees, how far three labellers agreed, and what it all cost. It also says what the benchmark does not cover, and one thing it found that this change does not fix: a step that says nothing is never judged. In click-zh only 8 and 10 of 19 moves between issues came with a step taskcut would judge, and make eval-mechanism now fails one check twice out of two, on a turn of four tasks the model did without a word. The README's limitations and troubleshooting say so. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
$.session.messages()already holds the step when it is judged, a block at a time; 0.8 left out only a last message matching the step whole, which never happens./taskcutand exact accounting of what the judge spends (4e7d40f).$.model.completeis not on the session's cost ledger, so/costnever showed the judge. taskcut now sums the API's own token counts in memory (free, no setting, nothing written), shows them in/taskcut, and ends each judgement's debug line with[judge sonnet: in=… cache_read=… cache_write=… out=… ms=…]. Claude Code's--debugis the debug mode; no taskcut-specific one.floorPercent: 0now judges from the first step (it waited for 5%).make eval-replay(the judge over every step real sessions took, labelled, with intervals and prices;REF=for another commit's judge;replay-comparefor paired differences),make eval-replay-check(the replay against a live session's prompts),replay-agree(label agreement), theclick-zhworkload (20 real click changes in one Chinese message, task list, a commit per issue), and the judge's cost in every end-to-end report.Why
The judge's prompt was nine tenths commands in full and grew with the session: three cents a judgement, $2.57 over the stretch a 35% floor judges on a long run. And "is it no worse?" had no instrument that could answer it.
How it was tested
Claude Code 2.1.280, Sonnet 5. Full numbers and caveats:
docs/measurement.md, "What the judge reads, replayed over real sessions".The judge, 228 labelled steps from 7 committed sessions (82 boundaries, 129 not), 3 repeats, 95% intervals:
Paired: boundaries +0.0% (−2.4 to +2.4), false NEXT −3.9% (−7.5 to −0.8). With one unpublished dev session added the false-NEXT interval is −6.2 to +0.2.
How the instrument itself was checked:
eval/replay/disagreements.md. Please skim that file — the annotators share the judge's model family../scripts/validate.sh(79 tests, tsc against regenerated 2.1.280 declarations,claude plugin validate).End to end: sqlglot-long
default— 36/36 issues, every check, 1 judgement (1,952 in / 47 out), 1 compaction. click-zh off/on — 23/23 checks each (never reached the floor)./taskcutchecked in a real session against the debug log: identical../scripts/validate.shpassestsc -p tsconfig.jsonpasses against regenerated declarations (2.1.280)Known gap this PR does not fix
A step that says nothing is never judged. In click-zh only 8 and 10 of 19 issue-to-issue moves had a judgeable step (the rest: a silent commit + the next issue's first command). And
make eval-mechanismfails 1 of 9 checks in 2 of 2 runs today: the model did the second four-task turn without a word, so that turn got no compaction. Every judgement that was made was right. Likely fix for a follow-up: treat a step that updates the task list (task k completed, k+1 in progress) or commits as judgeable, and say so in the judge's rules — which needs new labels for those steps, so it is its own change.Spend
About $49 at list price for this round's runs ($32 replay, of which $26 is the old judge; $9 sqlglot; $5 click-zh; the rest checks).
🤖 Generated with Claude Code