Skip to content

feat!: compact when the work moves on, not when a piece is finished - #5

Merged
wasd96040501 merged 1 commit into
mainfrom
feat/compact-before-next-piece
Sep 23, 2026
Merged

wasd96040501 merged 1 commit into
mainfrom
feat/compact-before-next-piece

Conversation

@wasd96040501

@wasd96040501 wasd96040501 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

What this changes

taskcut now compacts where the work moves on from a finished piece to another, not where a piece is finished. Hand over A, B and C in one message: it compacts between A and B and between B and C, and no longer after C. The end of a turn is no longer judged or compacted: the work is back with the person, and /compact before new work is theirs. Otherwise the structure is 0.8's: steps judged inside a turn, the turn ended, compacted, carried on with Continue..

One commit, on top of #6 (the $.model.complete reply-shape fix, merged): the question changes from "finished?" (DONE/WORKING) to "moves on?" (NEXT/SAME); the end of a turn is no longer judged; no step is judged when it says nothing, after a compaction until the model has changed something, or at a turn's first step (see below); the judge reads the whole conversation, as the permission classifier does. make eval-judge is added: labelled steps through the same call.

Why

  • Compacting after the last piece helps nothing: nothing follows yet, and the person's next message is often about that piece ("also handle None", "commit that").
  • A Claude Code 2.1.280 bug: a compaction after a request that ended on the person's words, /compact typed by hand included, answers them instead of summarising. taskcut ends a turn only after a step that followed tool results. 0.8 was exposed to it at every turn answered without a tool.
  • The judge's prompt, per the classifier's source (2.1.280 bundle, xLt/Zke/iRr): the classifier reads its whole transcript and never leaves the oldest out; too long is no verdict. It is cached because it builds its own request with four cache_control breakpoints. $.model.complete (YTe) sends none, and the usage it reports is the API's own, so the judge is uncached today — no plugin-side fix without copying engine internals. The judge now reads the whole conversation too, so each prompt begins with everything the one before it read, ready for a call that caches.

How it was tested

steps 0.8 question this PR
a piece done, another asked for 18/18 18/18
the last piece done 0/15 15/15
a piece in progress 12/15 15/15

(make eval-judge, 16 steps × 3, sonnet.) make eval-mechanism in a real interactive session, two four-task messages and a recall between: 9/9 checks, six compactions all between pieces, none after the last piece or at a turn's end, every file right.

Judge prompt, replaying 104 real steps: 15k characters on average, 30.5k at most; each prompt begins with the whole of the one before up to its step.

  • make validate passes
  • tsc passes against regenerated declarations
  • Ran in a real interactive session (make eval-mechanism, eval driver on a pty)
  • Claude Code version tested against: 2.1.280 (2.1.278's reply shape covered by unit tests only)

https://claude.ai/code/session_01Hy8i9XZbHT6uPuVZwAP8vT

Hand over A, B and C in one message: taskcut compacted after C too, where
nothing follows yet and the person's next message is often about C. The
judge now asks whether the work moves on from a finished piece to another
(NEXT/SAME) rather than whether a piece is finished (DONE/WORKING), and the
end of a turn is no longer judged or compacted: the work is back with the
person, and /compact before new work is theirs. The structure is otherwise
0.8's: steps judged inside a turn, the turn ended, compacted, carried on.

A step is not judged when it says nothing, nor after a compaction until the
model has changed something (the summary made the next piece's first step
look like the move to it), nor at a turn's first step: on Claude Code
2.1.280 a compaction after a request that ended on the person's words,
/compact typed by hand included, answers them instead of summarising.

The judge reads the whole conversation, as auto mode's permission classifier
reads its whole transcript (2.1.280, xLt/Zke), rather than the newest 40,000
characters; too long for its window is no verdict. Each prompt begins with
everything the one before it read. $.model.complete sends no cache_control
(YTe), so it is uncached today; design.md says what a cached call needs.

make eval-judge: 16 labelled steps x3, 48/48 (0.8's question 30/48).
make eval-mechanism: two four-task messages and a recall, 9/9.

Claude-Session: https://claude.ai/code/session_01Hy8i9XZbHT6uPuVZwAP8vT
@wasd96040501
wasd96040501 force-pushed the feat/compact-before-next-piece branch from 342a976 to 9d448de Compare September 23, 2026 01:07
@wasd96040501
wasd96040501 merged commit 311ffa8 into main Sep 23, 2026
3 checks passed
@wasd96040501
wasd96040501 deleted the feat/compact-before-next-piece branch September 23, 2026 01:17
@wasd96040501 wasd96040501 mentioned this pull request Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant