Repository navigation
feat!: compact when the work moves on, not when a piece is finished - #5
Merged
Merged
Conversation
3 tasks done
Hand over A, B and C in one message: taskcut compacted after C too, where nothing follows yet and the person's next message is often about C. The judge now asks whether the work moves on from a finished piece to another (NEXT/SAME) rather than whether a piece is finished (DONE/WORKING), and the end of a turn is no longer judged or compacted: the work is back with the person, and /compact before new work is theirs. The structure is otherwise 0.8's: steps judged inside a turn, the turn ended, compacted, carried on. A step is not judged when it says nothing, nor after a compaction until the model has changed something (the summary made the next piece's first step look like the move to it), nor at a turn's first step: on Claude Code 2.1.280 a compaction after a request that ended on the person's words, /compact typed by hand included, answers them instead of summarising. The judge reads the whole conversation, as auto mode's permission classifier reads its whole transcript (2.1.280, xLt/Zke), rather than the newest 40,000 characters; too long for its window is no verdict. Each prompt begins with everything the one before it read. $.model.complete sends no cache_control (YTe), so it is uncached today; design.md says what a cached call needs. make eval-judge: 16 labelled steps x3, 48/48 (0.8's question 30/48). make eval-mechanism: two four-task messages and a recall, 9/9. Claude-Session: https://claude.ai/code/session_01Hy8i9XZbHT6uPuVZwAP8vT
wasd96040501
force-pushed
the
feat/compact-before-next-piece
branch
from
September 23, 2026 01:07
342a976 to
9d448de
Compare
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
taskcut now compacts where the work moves on from a finished piece to another, not where a piece is finished. Hand over A, B and C in one message: it compacts between A and B and between B and C, and no longer after C. The end of a turn is no longer judged or compacted: the work is back with the person, and
/compactbefore new work is theirs. Otherwise the structure is 0.8's: steps judged inside a turn, the turn ended, compacted, carried on withContinue..One commit, on top of #6 (the
$.model.completereply-shape fix, merged): the question changes from "finished?" (DONE/WORKING) to "moves on?" (NEXT/SAME); the end of a turn is no longer judged; no step is judged when it says nothing, after a compaction until the model has changed something, or at a turn's first step (see below); the judge reads the whole conversation, as the permission classifier does.make eval-judgeis added: labelled steps through the same call.Why
/compacttyped by hand included, answers them instead of summarising. taskcut ends a turn only after a step that followed tool results. 0.8 was exposed to it at every turn answered without a tool.xLt/Zke/iRr): the classifier reads its whole transcript and never leaves the oldest out; too long is no verdict. It is cached because it builds its own request with fourcache_controlbreakpoints.$.model.complete(YTe) sends none, and the usage it reports is the API's own, so the judge is uncached today — no plugin-side fix without copying engine internals. The judge now reads the whole conversation too, so each prompt begins with everything the one before it read, ready for a call that caches.How it was tested
(
make eval-judge, 16 steps × 3, sonnet.)make eval-mechanismin a real interactive session, two four-task messages and a recall between: 9/9 checks, six compactions all between pieces, none after the last piece or at a turn's end, every file right.Judge prompt, replaying 104 real steps: 15k characters on average, 30.5k at most; each prompt begins with the whole of the one before up to its step.
make validatepassestscpasses against regenerated declarationsmake eval-mechanism, eval driver on a pty)https://claude.ai/code/session_01Hy8i9XZbHT6uPuVZwAP8vT