Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
{
"$schema": "https://anthropic.com/claude-code/marketplace.schema.json",
"name": "taskcut",
"description": "Compaction at the end of a sub-task, not when the window fills, for Claude Code",
"description": "Compaction when the work moves on to the next sub-task, not when the window fills, for Claude Code",
"owner": {
"name": "Yangze Luo"
},
"plugins": [
{
"name": "taskcut",
"description": "Compacts at the end of a sub-task instead of when the context window fills. Once the context is past a floor, a small model judges at the end of each turn whether the work asked for is finished, from the same inputs auto mode's permission classifier reads; if it is, taskcut runs Claude Code's own compaction. Below the floor it does nothing and costs nothing.",
"description": "Compacts when the work moves on from one sub-task to the next instead of when the context window fills, including in the middle of a long turn. Once the context is past a floor, a model judges each step, reading what you asked for and where Claude has got to but never any tool output, for whether the work moves on from a finished piece to another; if it does, taskcut runs Claude Code's own compaction. Below the floor it does nothing and costs nothing.",
"author": {
"name": "Yangze Luo"
},
Expand Down
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "taskcut",
"version": "0.8.0",
"description": "Compacts when the work moves on from one sub-task to the next instead of when the context window fills, including in the middle of a long turn. Once the context is past a floor, a model judges each step, from the same inputs auto mode's permission classifier reads, for whether the work moves on from a finished piece to another; if it does, taskcut runs Claude Code's own compaction. Below the floor it does nothing and costs nothing.",
"description": "Compacts when the work moves on from one sub-task to the next instead of when the context window fills, including in the middle of a long turn. Once the context is past a floor, a model judges each step, reading what you asked for and where Claude has got to but never any tool output, for whether the work moves on from a finished piece to another; if it does, taskcut runs Claude Code's own compaction. Below the floor it does nothing and costs nothing.",
"author": {
"name": "Yangze Luo"
},
Expand Down
45 changes: 39 additions & 6 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,26 +23,59 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- A turn is never ended at its first step: on Claude Code 2.1.280 a compaction
after a request that ended on your own words, `/compact` typed by hand
included, answers them instead of summarising the conversation.
- The judge reads the whole conversation, as the permission classifier reads
its whole transcript, rather than the newest 40,000 characters, so each
judgement's prompt begins with everything the one before it read. A
conversation too long for the judge's window is no verdict, and keeps the
context.
- **The judge reads what says whether the work moves on, and little else**:
every message you sent, the first and the latest three whole and the rest
cut to a line; Claude's latest ten messages; what its latest twelve calls
touched -- a file, or what a command says it does, never the call in full;
its task list; and the step. It no longer reads every command in full, nor
`CLAUDE.md`, and taskcut no longer reads any file. A judgement is about two
thousand tokens however long the session has run, where it was sixteen
thousand on average and grew with the work, and costs about half a cent
instead of about three.
- The notice reads `compacting before the next piece`.
- taskcut is released under the MIT License, in place of Apache-2.0.

### Added

- `make eval-judge`: the judge alone, over labelled steps, three times each,
through the call taskcut makes.
- **`/taskcut`** says what taskcut has done in this session and what it cost:
how many steps it judged, the tokens those calls read and wrote, and the
compactions it started. The judge's calls are not in `/cost`, which counts
only the session's own requests, nor in the transcript; this is where they
are counted. The tally is the API's own token counts, summed in memory:
nothing is written anywhere and nothing is asked of a model.
- `make eval-replay`: the judge over every step taskcut would have judged in
real sessions -- 228 steps from seven, committed with their labels -- with
intervals on every rate, the list price of each judgement, and
`REF=` to measure the judge of another commit on the same steps.
`make eval-replay-check` checks, against a live session, that the replay
gives the judge the prompt taskcut builds there.
- The `click-zh` workload: the twenty click changes handed over in one
message in Chinese, with a task list and a commit per issue.
- Each judgement's debug line (`claude --debug`) ends with what the call cost,
`[judge sonnet: in=1834 cache_read=0 cache_write=0 out=52 ms=2140]`, and the
benchmark adds them up: its cost table now has the judge beside the session.

### Fixed

- **Nothing was compacted on Claude Code 2.1.280.** `$.model.complete` now
resolves `{ isAnswered, text, usage }` rather than the reply's text; taskcut
read the object as text, every judgement failed, and the failure was logged
only to the debug log. The reply is now read in either shape, so 2.1.278
keeps working, and a failed call is logged with its reason.
keeps working, and a failed call is logged with its reason -- for an API
error, its status and kind, so a spent rate limit reads as one.
- **The judge read the step it was judging twice**, once as the step and once
as the latest thing the assistant said, on most steps -- and whether it did
depended on how fast the step's tools ran. When a step is judged,
`$.session.messages()` already holds it, a block at a time (its words, then
each call, then any result already in), and only a last message matching
the step whole was left out. The step's messages are now found and left
out whatever their shape. Found by checking the replay benchmark against
the prompts a live session builds.
- **A `floorPercent` of 0 judged nothing until the context reached 5%.** The
wait taskcut keeps after a compaction that could not get under the floor
read "no compaction yet" as one that had left the context at 0%.

## [0.8.0] - 2026-09-22

Expand Down
23 changes: 23 additions & 0 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,29 @@ eval-mechanism: ## Check the mechanism end to end in one short real session (MOD
eval-judge: ## Ask the judge about the labelled steps, three times each (MODEL=)
@$(PYTHON) -m taskcut_eval.cli --results $(RESULTS) judge --model $(MODEL)

REF ?=
REPEATS ?= 3

.PHONY: eval-replay
eval-replay: ## The judge over every labelled step of real sessions (MODEL=, REF= a commit's judge, REPEATS=)
@$(PYTHON) -m taskcut_eval.cli --results $(RESULTS) replay --model $(MODEL) --repeats $(REPEATS) $(if $(REF),--ref $(REF))

.PHONY: eval-replay-compare
eval-replay-compare: ## Two replay runs over the same steps, B minus A (A=, B= result files)
@$(PYTHON) -m taskcut_eval.cli replay-compare $(A) $(B)

.PHONY: eval-replay-check
eval-replay-check: ## Check the replay builds the prompts a live session gives the judge (MODEL=)
@$(PYTHON) -m taskcut_eval.cli replay-check --model $(MODEL)

.PHONY: eval-replay-build
eval-replay-build: ## A replay set from a transcript, with a labels file to fill (TRANSCRIPT=, NAME=, SOURCE=)
@$(PYTHON) -m taskcut_eval.cli replay-build --transcript $(TRANSCRIPT) --name $(NAME) --source "$(SOURCE)"

.PHONY: eval-replay-sheet
eval-replay-sheet: ## Every judged step of a set with its context, for labelling (SET=)
@$(PYTHON) -m taskcut_eval.cli replay-sheet --set $(SET) --unlabelled

.PHONY: eval-report
eval-report: ## Render the collected transcripts
@$(PYTHON) -m taskcut_eval.cli --results $(RESULTS) report
Expand Down
44 changes: 28 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,15 +23,15 @@ when the context window fills.

taskcut changes **when** Claude Code compacts, not how. Once the context is
past a floor, a model judges each step Claude takes for whether the work is
moving on from a finished piece to another, reading what auto mode's
permission classifier reads. If it is, taskcut runs Claude Code's own
compaction, the one `/compact` runs. A piece being finished is not enough:
when the last thing you asked for is done, nothing has moved on yet, and what
you say next may well be about it. Hand over twenty tasks in one message and
walk away, and it compacts between them, not after the last. Once a turn has
ended the work is back with you, and so is `/compact`. Below the floor it does
nothing at all and costs nothing, and nothing in your prompts has to mention
it.
moving on from a finished piece to another, reading what you asked for and
where Claude has got to, never any tool output. If it is, taskcut runs Claude
Code's own compaction, the one `/compact` runs. A piece being finished is not
enough: when the last thing you asked for is done, nothing has moved on yet,
and what you say next may well be about it. Hand over twenty tasks in one
message and walk away, and it compacts between them, not after the last. Once
a turn has ended the work is back with you, and so is `/compact`. Below the
floor it does nothing at all and costs nothing, and nothing in your prompts
has to mention it.

## Try it

Expand Down Expand Up @@ -114,12 +114,16 @@ Start sessions with `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude`, or put
| | |
| --- | --- |
| **Below the floor** | Nothing: no model call, no tool, nothing written. |
| **Each step Claude explains past it** | One `sonnet` call and a sentence out. What goes in is your messages and Claude's commands since the last compaction, never their output: a few thousand tokens, tens of thousands in a long stretch of work, uncached (`$.model.complete` marks no cache point) — a few cents. It runs while the step's tools run, so it rarely adds a wait. |
| **Each step Claude explains past it** | One `sonnet` call and a sentence out. What goes in is your messages, Claude's latest messages, what its latest commands touched and the step it is taking, never any output: about two thousand tokens however long the session has run, uncached (`$.model.complete` marks no cache point) — about half a cent. It runs while the step's tools run, so it rarely adds a wait. |
| **Each compaction** | Whatever `/compact` costs, because it is `/compact`. |

Past the floor is a short stretch: a compaction takes the context back under
it, and judging stops until it fills again.

`/taskcut` says what it has spent so far in the session. `/cost` does not
count the judge's calls -- Claude Code keeps no ledger of a plugin's model
calls -- so they are counted there, from the token counts the API returns.

Thirty-six real sqlglot changes, handed to Sonnet 5 in one message and left
to run: taskcut compacted once inside the turn, after issue 20, from 431,567
tokens to 10,511, and carried on. The peak context fell from 71% of the window
Expand All @@ -145,12 +149,14 @@ installed with.
After each step of the main conversation, taskcut reads the context fill the
status line shows. Below `floorPercent` it stops there. Past it, it asks
`model`, through the hooks API's `$.model.complete`, whether the work moves on
from a finished piece to another at that step. The judge reads what auto mode's
permission classifier reads — your messages, the assistant's tool calls except
read-only lookups, and `CLAUDE.md`, never any tool output, all of it however
long the session — plus the step it is judging: what Claude just said and the
calls it is making. A step that asks you something is never a boundary, and
the end of a turn is never judged.
from a finished piece to another at that step. The judge reads what says so
and little else: every message you sent (the first and the latest few whole,
the rest cut to a line), Claude's latest messages, what its latest calls
touched — a file, or what a command says it does, never the call in full — its
task list if it keeps one, and the step it is judging: what Claude just said
and the calls it is making. Never any tool output, and not `CLAUDE.md`. A step
that asks you something is never a boundary, and the end of a turn is never
judged.

On a yes, taskcut waits for that step's tools to finish, ends the turn with
`$.turn.abort` before the next request goes out, calls `$.session.compact()` —
Expand All @@ -168,6 +174,12 @@ reasoning, and the designs this one replaced.
compact, and taskcut never ends a turn there.
* **The judge never sees tool output**, so a reply that claims more than was
done can fool it. The cost is a compaction a little early.
* **Only a step that says something is judged.** A step that finishes one
piece and starts the next without a word -- a commit, then the next issue's
first command -- is never asked about. Claude usually says so somewhere
nearby, but not always: in the benchmark about half the moves between issues
were silent, and one turn of four tasks done in silence was not compacted at
all ([measurement](docs/measurement.md)).
* **A compaction inside a turn splits it in two.** Claude Code compacts only
between turns, so taskcut ends the turn and starts the next with
`Continue.`, which the transcript shows as a message from the plugin. What
Expand Down
2 changes: 1 addition & 1 deletion SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ this plugin calls. Reports that matter most:
session, or in a subagent's loop;
* a prompt taskcut submits that is anything other than `Continue.`;
* a path by which the judge is sent more than it is documented to read -- tool
output, or files other than the `CLAUDE.md` files already in the context.
output, a call in full from before the step, or the contents of any file.

## What is out of scope

Expand Down
8 changes: 5 additions & 3 deletions docs/compatibility.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,12 +8,14 @@ declarations it is written against are generated per Claude Code version by

| Depends on | Kind | If it changes |
| --- | --- | --- |
| `session.start`, `turn.step`, `turn.complete` | hook events | taskcut stops working; `claude plugin validate` reports the unknown event before a session loads it |
| `$.session.usage`, `$.session.messages`, `$.session.compact`, `$.fs.read`, `$.env.get`, `$.model.complete`, `$.turn.abort`, `$.prompt.submit`, `$.ui.log` | engine calls | same: refused at load, named in the validation output |
| `session.start`, `turn.step`, `turn.complete`, `command.run` | hook events | taskcut stops working; `claude plugin validate` reports the unknown event before a session loads it |
| `$.session.usage`, `$.session.messages`, `$.session.compact`, `$.env.get`, `$.model.complete`, `$.turn.abort`, `$.prompt.submit`, `$.ui.log`, `$.command.register` | engine calls | same: refused at load, named in the validation output. `/taskcut` failing to register is logged to the debug log and changes nothing else |
| `$.model.complete`'s `usage`, and the call being off the session's cost ledger | data shape, engine behaviour, 2.1.280 | `/taskcut` counts calls whose usage it cannot read apart, as unmetered. If the ledger starts counting them, `/cost` would include what `/taskcut` also reports |
| what `$.session.messages` holds: a compaction's summary opening with "This session is being continued from a previous conversation", a plugin's prompt with "The … plugin sent a message", a background task's report with `<task-notification>` | text Claude Code writes | the judge would read such a message as something you asked for; it is still shown the step, so the cost is a request misread, not a crash |
| a hook's budget counting only its own code, not its `$` calls | engine rule | the judgement inside a step would overrun it, and the engine would skip the hook: taskcut would compact only at the end of a turn, and say so in the transcript |
| what `$.model.complete` resolves to | data shape | `claude plugin validate` does not see it; only `tsc` against regenerated types does. 2.1.278, which taskcut was measured on, resolved the reply's text; 2.1.280 resolves `{ isAnswered, text, usage }`, and 0.8.0 read that object as text, so every judgement failed and nothing was ever compacted. The reply is now taken as `unknown` and read by `readReply`, which knows both shapes and treats anything else as no reply |
| a compaction after a request that ended on the person's words answering them instead of summarising | engine bug, 2.1.280 | taskcut never ends a turn at its first step; when it is fixed, that guard can go |
| `$.model.complete` sending no cache breakpoint | engine behaviour, 2.1.280 | none needed: the judge's prompt keeps a stable prefix, so a release that caches it serves it with no change here |
| `$.model.complete` sending no cache breakpoint | engine behaviour, 2.1.280 | little: the judge reads about two thousand tokens, and a cache would save a fraction of a cent; see [the judge's prompt cache](design.md#the-judges-prompt-cache) |
| `session.start`'s `isInteractive` | data shape | taskcut would stop ending turns early, and so would not compact at all |
| `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS` | early-access flag | **nothing**, by design — see below |

Expand Down
Loading
Loading