Quick start · How it works · Troubleshooting · Limitations · FAQ · Blueprint
Your coding agent uses its most expensive model for every small decision. It sends every test log, build log, and search result to that model. Then it carries the output into every later turn, and you pay for it again each time.
compactio adds a System 1. After each tool call, a fast decision model (Jev) answers one question: how much of this output does the agent need for the current goal? Jev does not write text. It returns a typed choice in about half a second, and code applies that choice. The LLM reads only what it needs.
A real Claude Code run: 20,501 → 747 characters (−96%) in 0.75 s for $0.00005. The agent still found the cause.
In Thinking, Fast and Slow, Daniel Kahneman describes two systems. System 1 is fast and automatic. System 2 is slow and deliberate. Coding agents have only System 2.
| ~50% | Share of a Claude Code context that is tool output (ctxlens profile) |
| 97% | Share of billed tokens that are cache reads: old context, paid again on each turn (devxlabs) |
| Late | Compaction starts only near the limit, costs one more large request, and can lose instructions |
compactio stops the waste where it starts: at the tool output, before it enters the context.
- Decides, does not generate. Jev picks one of four views:
full,errors,headtail,stub. - Safe by default. Low confidence, a timeout, or an API error returns the full output. compactio never blocks your agent.
- Nothing is lost. Every cut output stays on your disk. The note in the output tells the agent how to get it back.
- Code is never cut. A
Readof a source file always passes in full. compactio only skips an exact re-read of an unchanged file. - Data files are cut. A
Readof a log, a CSV, a JSONL dump, a lockfile, or a minified bundle takes the same path asBashoutput. - Old output can go too (opt-in). The Sweep proxy removes old tool results that the current goal no longer needs. See Sweep.
- Secrets stay local. API keys, tokens, private keys, and
KEY=valuepairs are masked before any request. - Two providers. Use a TypeSafe key or an OpenRouter key.
- Works without a key. Local mode uses lossless filters, the re-read skip, and a head-and-tail cut for very large output.
- Zero dependencies. No build step for the plugin, and no
node_modules.
compactio has two parts:
| Part | What it does | Where it runs | Setup |
|---|---|---|---|
| Filter | Cuts large new tool output before the agent reads it | Every Claude Code session: terminal and desktop app | Install the plugin |
| Sweep (optional) | Removes old tool output that the goal no longer needs | claude in a terminal only |
One more command |
Node 22.18 or later must be on your PATH. Check with node -v.
Run this in a terminal:
claude plugin marketplace add RamaAditya49/compactio
claude plugin install compactio@compactioRestart Claude Code. The filter now works in local mode, with no key.
With a key, Jev decides how much of each output to keep. Without a key, compactio cuts only very large output, with a fixed rule.
-
Get a key from TypeSafe, the maker of Jev, or from OpenRouter.
-
Run this in a terminal, and paste the key when it asks. The key does not show on the screen.
npx compactio key
-
Restart Claude Code.
Do not paste the key into the Claude Code chat. Text in the chat goes into the conversation.
Use the Sweep if you run long sessions with claude in a terminal. The Claude desktop app does not use it (see Limitations).
In Claude Code, run:
/compactio:setup
Setup checks the key and turns on the Sweep. Then restart your claude sessions. Details: Sweep.
/compactio:gain
compactio · all sessions
────────────────────────────────────────────────────
Tokens kept out of context ~12,478
Share of tool output cut ████░░░░░░░░░░░░░░░░ 20%
Outputs cut 3 of 17 (avg cut 93%)
Jev decisions 4 · $0.0002
Old outputs swept 0
Unchanged re-reads skipped 0
Fail-open 0
────────────────────────────────────────────────────
Biggest cuts
Bash headtail 30.0k → 408 chars (local)
Bash errors 20.5k → 747 chars (jev)
Bash clean 3.0k → 2.4k chars (lossless)
────────────────────────────────────────────────────
tokens ≈ characters ÷ 4
The numbers grow as you work. A short session with small outputs shows 0, and that is correct: small output passes untouched.
For the Sweep, run /compactio:sweep status. It must say proxy running and settings point at it.
| To turn off | Run |
|---|---|
| The Sweep | /compactio:sweep off, or npx compactio sweep off in a terminal. Then restart your claude sessions. |
| All of compactio | npx compactio sweep off, then claude plugin disable compactio@compactio |
Turn off the Sweep before you remove the plugin. Otherwise Claude Code still points at a proxy that is gone.
compactio uses the Claude Code PostToolUse hook. The hook runs after a tool finishes and before the model sees the result.
| Step | Who | What |
|---|---|---|
| 1 | Tool | Output arrives from Bash, Grep, WebFetch, WebSearch, an MCP tool, or Read. |
| 2 | Code | Output below 2,000 characters passes. An exact re-read of an unchanged file becomes one line. |
| 3 | Code | Remove ANSI codes, repeated lines, and blank runs. No information is lost. |
| 4 | Jev | For output above 6,000 characters, choose the view that the current goal needs. |
| 5 | Code | Apply the view, store the original on disk, and add a one-line note. |
| 6 | LLM | Read the result. |
| View | The agent keeps | Typical case |
|---|---|---|
full |
All of the output | The output is the answer, or the agent will edit from it |
errors |
Error, failure, and warning lines with context, plus the tail | A test run with one failure |
headtail |
The first 40 and the last 30 lines | A long build or install log |
stub |
One line | Output that is not related to the goal |
The hook can only cut new output. Old tool results stay in the history, and the agent sends them again on every turn. The Sweep removes them.
The Sweep is a local proxy between Claude Code and the Anthropic API. On each request it does three steps:
| Step | Who | What |
|---|---|---|
| 1 | Code | Put back every earlier tombstone, so that the prompt prefix does not change between turns. |
| 2 | Jev | When old results hold 40,000 characters or more, rate up to 16 of them in one request: keep or drop. |
| 3 | Code | Replace the drop results with a tombstone, but only when the saving pays for the prompt-cache rewrite. |
A tombstone looks like this:
[compactio: the output of Bash npm test was removed because the current goal no longer needs it. Full output: node ".../cli.js" show 1a2b3c4d. Or run the tool again.]
Rules:
- The last 10 messages are never swept. They are the work in progress.
- Results below 4,000 characters and results with images are never swept.
- Jev must answer
dropwith a confidence of 0.8 or more. Otherwise the result stays. - Cache gate. A change in the middle of the history makes the next request write the cache again after that point. The Sweep drops only when
dropped × 0.1 × turns ≥ rest-of-history × 1.15.turnsisCOMPACTIO_SWEEP_TURNS(default 30).
/compactio:setup turns it on when a Jev key is present. You can also use /compactio:sweep on|off|status, or npx compactio sweep on|off|status in a terminal.
What sweep on changes on your computer
- Copy compactio to
~/.compactio/bin, so that a plugin update does not break the proxy. - Start the proxy. On Linux, it is the systemd user service
compactio-proxy, withRestart=always: it starts again after a crash and after a reboot. On macOS and Windows, it is a background process. - Wait until the proxy answers.
- Add two keys to the
envblock of~/.claude/settings.json(backup:settings.json.compactio-bak):ANTHROPIC_BASE_URL=http://127.0.0.1:8787ANTHROPIC_DEFAULT_OPUS_MODEL=claude-opus-5-5[1m]. Behind a custom base URL, Claude Code does not detect the 1M window, and autocompact runs in a loop. The[1m]suffix fixes it.
Session guard. At the start of each Claude Code session, compactio checks the proxy. If the proxy is down, compactio starts it before the first request. If it still does not start, Claude Code shows a message with the command that turns the Sweep off.
compactio sweep status shows the state. compactio sweep off removes the keys first, then stops the proxy.
Only a bounded and redacted summary goes to the provider:
- the goal: your last prompt, up to 2,000 characters
- the tool name and its input, up to 600 characters
- a preview of the output: the first lines, the error lines, and the last lines, about 6,000 characters in total
The full output never leaves your machine. Without an API key, nothing leaves your machine.
Set these variables in the env block of ~/.claude/settings.json.
| Variable | Default | Description |
|---|---|---|
TYPESAFE_API_KEY |
— | Use Jev through TypeSafe. It has priority when both keys are set. |
OPENROUTER_API_KEY |
— | Use Jev through OpenRouter. Requests show as the app compactio on OpenRouter. |
COMPACTIO_JEV_MODEL |
jev-1.13.0 (TypeSafe), ~typesafe/jev-latest (OpenRouter) |
Jev model. |
COMPACTIO_JEV_URL |
the provider endpoint | Custom endpoint, for example a proxy. |
COMPACTIO_TIMEOUT_MS |
1500 |
Time limit for one decision. After it, the full output passes. |
COMPACTIO_HOME |
~/.compactio |
Folder for stored outputs, read hashes, Sweep decisions, and the decision log. |
COMPACTIO_UPSTREAM |
https://api.anthropic.com |
Sweep proxy: the API that receives the requests. |
COMPACTIO_PORT |
8787 |
Sweep proxy: the local port, when proxy gets no port. |
COMPACTIO_SWEEP_TIMEOUT_MS |
3000 |
Sweep proxy: time limit for one Jev request. After it, the history passes unchanged. |
COMPACTIO_SWEEP_TURNS |
30 |
Sweep proxy: expected turns left in a session. A higher value drops more often. |
| Command | Description |
|---|---|
/compactio:setup |
Check the Jev key and turn on the Sweep. |
npx compactio key |
Save a Jev key in the Claude Code settings. The key does not show on the screen. |
/compactio:gain |
Show the savings scoreboard in Claude Code. |
npx compactio gain |
Show the scoreboard in a terminal. |
npx compactio show <id> |
Print a stored original output. The agent runs this itself when it needs the full output. |
npx compactio sweep on|off|status or /compactio:sweep on|off|status |
Install, remove, or check the Sweep proxy service. |
npx compactio proxy [port] |
Run the Sweep proxy in the foreground on 127.0.0.1. |
claude plugin disable compactio@compactio |
Turn compactio off. |
| Problem | Cause | Fix |
|---|---|---|
Claude Code cannot connect to the API (ECONNREFUSED) |
Claude Code points at the Sweep proxy, and the proxy is down | Run npx compactio sweep off in a terminal, then restart Claude Code |
| "Autocompact is thrashing" | Claude Code does not know the model's 1M window behind the proxy | Pick the opus model, or add [1m] to the model name. See Limitations. |
/compactio:sweep status says the desktop app sets its own API address |
The desktop app does not use the Sweep | Use claude in a terminal for long sessions. The filter still works in the desktop app. |
/compactio:gain shows 0 |
No large output yet | Work as usual. Only output above 2,000 characters counts. |
npx compactio ... prints nothing |
Version 0.2.0 or earlier | Run npx compactio@latest ... |
| Jev decisions stay at 0 | No key, or Claude Code did not restart after compactio key |
Run /compactio:setup to check the key, then restart Claude Code |
Can compactio break my agent?
compactio fails open. If Jev is slow, returns an error, or has low confidence, the agent gets the full output. Source files from Read are never cut. Every cut output is stored, and the note in the output tells the agent how to get it back.
Does it send my code to a third party?
Only when you set an API key, and only a redacted preview of large tool outputs (see What leaves your machine). The redaction uses patterns, so it cannot find every possible secret. If that is a concern for a project, run compactio without a key.
How much does it cost?
One Jev decision uses about 1,000 input tokens. At $0.042 per million input tokens, that is about $0.00005. Output tokens are free. Most tool outputs are small and need no decision at all.
Does it replace compaction?
No. It makes compaction less frequent, because less output enters the context. Claude Code still compacts when the window is full.
Does it work with rtk?
Yes. rtk shrinks known shell commands before they run. compactio acts after each tool call and also covers Grep, web tools, MCP tools, and re-reads.
Which agents are supported?
Claude Code today. Codex CLI, OpenCode, Gemini CLI, Cursor, and Trae are next. See the roadmap.
- Token counts are estimates: characters ÷ 4.
- "Hundreds of times cheaper" applies to one decision (Jev compared with an LLM), not to your total bill.
- Claude Code already limits Bash output to 30,000 characters. On one Bash call, compactio saves at most about 7,500 tokens. The agent then carries that saving through every later turn.
- We make no claim about the total bill until the public benchmark (task success, tokens, and cost) exists.
Replay on real sessions (node scripts/replay.ts --jev). The 10 largest Claude Code sessions of the author, from before compactio, ran through the filter:
| Tokens | Share | |
|---|---|---|
| Tool output (text) | ~2.52M | 44% of context |
| Kept out by the filter, with Jev | ~44k | 1.7% of tool output, 0.8% of context |
| Kept out by the filter, local mode | ~29k | 1.2% of tool output |
Why the filter saves little in these sessions:
- 43% of the tool text came from outputs below 2,000 characters. The filter does not touch them.
- 24% came from outputs of 2,000 to 6,000 characters. Only the lossless clean runs on them.
- For most larger outputs, Jev answered
full. compactio keeps the full output when Jev is not sure.
The Sweep acts on old output instead. In one real test session, the goal changed, and the Sweep dropped 3 old reads: 98,031 → 879 characters. That saving repeats on every later turn. It is one session, not a benchmark.
Read these before you use compactio. They are the limits of the design, not bugs.
Where compactio has no effect
- Your messages, the agent's replies, the system prompt, and MCP tool schemas. compactio only touches tool output.
- Source code from
Read. It always passes in full, because the agent may edit it. In a session that mostly reads code, the saving is small. - Images, PDFs, and notebooks from
Read. They pass untouched and do not count in the scoreboard. - Old context, without the Sweep. The hook only cuts new output. The history that is already in the context stays until you run the Sweep,
/compact, or/clear. - Other hosts. compactio works in Claude Code only.
Limits of the decision
- Jev sees a preview, not the full output. The preview is the head, the error lines, and the tail, about 6,000 characters (2,400 for each Sweep result). Jev can misjudge an output whose important part is in the middle.
- Jev request limits. State and questions must fit in 64,000 tokens, and state plus the longest question in 32,000 tokens. For this reason, one Sweep request rates 16 results at most. The others wait for a later request.
- The goal is the last prompt. A short prompt such as "continue" gives Jev little to work with.
- The data-file list is fixed. compactio knows a data file by its name:
.log,.out,.csv,.tsv,.jsonl,.ndjson,.map,.min.js,.min.css, and common lockfiles. If the agent must edit such a file, it gets a cut view. It must runcompactio show <id>to see all of it. - Latency. A large output waits up to 1.5 s for Jev. A Sweep request waits up to 3 s. The time limit then lets the output pass unchanged.
Limits of the Sweep proxy
- Claude Code cannot reach the API when the proxy is down. The proxy fails open for its own errors, but not for a stopped process. The session guard starts it at the start of a session, and systemd starts it after a crash on Linux. A proxy that stops in the middle of a session on macOS or Windows stays down until the next session. Use
compactio sweep off, notsystemctl stop, to turn it off. - The Claude desktop app does not use the Sweep. It sets its own API address for its sessions, over
~/.claude/settings.json. The Sweep runs inclaudesessions in a terminal. The filter runs everywhere. - macOS and Windows are not tested. The background-process path is tested on Linux only.
- The key step needs a terminal. A key typed into the Claude Code chat goes into the conversation, so
/compactio:setupdoes not take a key. - Each sweep costs one cache rewrite. The cache gate estimates the cost with a fixed number of turns left (
COMPACTIO_SWEEP_TURNS). If the session ends sooner, the sweep costs more than it saves. - A tombstone is permanent. A dropped result stays dropped for the whole session. The agent must run the tool again or run
compactio show <id>. - The proxy sees all API traffic, including the auth header. It forwards the header and does not store it. It stores the dropped outputs on disk under
COMPACTIO_HOME. - The context window is set by name. Behind a custom
ANTHROPIC_BASE_URL, Claude Code does not detect the 1M window.sweep onmaps theopusalias toclaude-opus-5-5[1m]. If you pick a model by its full id, or pick Sonnet or Haiku, add[1m]yourself where the model supports it. When a new Opus ships, updateANTHROPIC_DEFAULT_OPUS_MODEL.CLAUDE_CODE_MAX_CONTEXT_TOKENSandCLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENTdid not stop the loop in Claude Code 2.1.282. - Test coverage. Unit tests use a fake API and a fake Jev. Real Claude Code sessions (Opus 5.5, subscription login) ran through the proxy: Jev rated old reads in 0.7 s, and the session guard started a stopped proxy. A real session where Jev drops a result is not tested yet.
Limits of the numbers
- Token counts are estimates: characters ÷ 4.
- The state files (
store/,sessions/,sweep.json) are never pruned.
| compactio | rtk | fast-jev-compaction | |
|---|---|---|---|
| When it acts | After each tool call | Before each shell command | At compaction |
| How it decides | Jev, from the current goal and a content preview | Fixed rules per command | Jev, from a size note |
| Tools covered | Bash, Grep, Web, MCP, data-file reads, re-reads, old results (Sweep) | Shell commands | All, at compaction |
- v0.1 Claude Code: tool output filter, re-read skip, scoreboard, redaction, local mode, TypeSafe and OpenRouter
- v0.2 Sweep: remove stale context in long sessions (opt-in local proxy), with prompt-cache protection. Data-file reads. One-step setup and session guard.
- v0.3 Codex CLI, OpenCode, Gemini CLI, Cursor, Trae. Replay evaluation on real sessions.
- v0.4 Gate: route each prompt to the cheapest model that can do the task
- v1.0 Public benchmark
See BLUEPRINT.md for the full design.
git clone https://github.com/RamaAditya49/compactio.git
cd compactio
node --test test/*.test.ts # tests
node scripts/assets.mjs # regenerate the README images
node scripts/replay.ts --jev # replay your own Claude Code sessions through the filterUse your local copy in Claude Code:
claude plugin marketplace add ./
claude plugin install compactio@compactioSee CONTRIBUTING.md for guidelines. Report security issues as described in SECURITY.md.
- Jev by TypeSafe AI, the decision model that powers compactio.
- Thinking, Fast and Slow by Daniel Kahneman, for the System 1 and System 2 model.
- rtk and fast-jev-compaction, which showed that this problem matters.
compactio is not affiliated with TypeSafe AI, OpenRouter, or Anthropic.
MIT © 2026 Rama Aditya