Your agents already generate the data needed to improve their runtime. RunTune closes the loop.
RunTune is a runtime learning loop for coding agents. It turns real Claude Code, Codex, Cursor, Antigravity and model-router runs into governed changes to the system around them.
- Success → capability. Working patterns that repeat become skills, CLI paths, specialized sub-agents, or better routes.
- Failure → constraint. Failures that repeat become guard rules and routing checks.
- Change → measurement. Every adopted change is measured against a control, then kept, reviewed, or retired.
It tunes the harness, not the model. RunTune doesn't teach an agent to remember yesterday. It changes tomorrow's runtime based on what happened yesterday.
No weight training, and no autonomous self-modification: RunTune proposes evidence-backed changes, and a human decides what enters the runtime.
RunTune is the feedback loop of an internal platform for GTM engineers. It reads what the agents actually did, proposes the narrowest change that would fix a repeated failure or capture a repeated success, and measures each approved change against a control. It was developed against about 238,000 recorded tool calls and model requests from one team's production agents, and its evidence doc keeps the five recommendations that changed once checked against production.
A coding agent is a model inside a harness: the skills it can load, the sub-agents it can spawn, the guards on its tools, and the routes its model calls take. The model is someone else's to train; the harness is yours to tune.
This is the difference from agent memory: the next agent runs in a different environment, whether or not it recalls anything.
Most agent observability ends at a dashboard: runs → traces → a person reads the dashboard. Traces tell you what happened. RunTune asks what the runtime should learn from it, turns the answer into a gated, approved change, and then checks on future runs whether the change worked.
pip install git+https://github.com/kkrlstrm/runtune # zero dependencies
runtune demo # the whole loop on a synthetic trace, in a temp dir
runtune record claude # backfill your Claude Code history (~/.claude/projects)
runtune record codex # and/or Codex (~/.codex/sessions)
runtune record cursor # and/or Cursor's agent (its local state.vscdb)
runtune record antigravity # and/or Antigravity (~/.gemini/antigravity/brain)
runtune notify # one digest: a card + numbered proposals, with a desktop notification
runtune show --open # read it
runtune reply 1,3 # approve items 1 and 3; the rest are snoozed for 28 daysThis uses only the transcripts already on your machine: no accounts, no database, no network.
To keep it running, runtune install prints the hook for Claude Code, Codex, Cursor and Antigravity (you
paste it, because RunTune never edits host settings), and runtune schedule --install sends a
weekly digest, reading Cursor's store first when Cursor is installed.
RunTune can stay on your laptop, or run continuously around a deployed agent system. On a small VM it reads evidence from a shared telemetry database that your agents' recorders write to (warehouse sources), snapshots the model provider's bill daily, and derives and reviews weekly. It sends each digest to Slack or email, and every hour it checks for a decision. The first valid answer on any channel wins, and the other channels are told it was handled. An approved change becomes a branch and a pull request against the agent's harness, and measurement starts only once that change is merged.
AGENT RUNS
↓
EVIDENCE
↓
RUNTUNE ─────→ Slack / email
↑ ↓
│ human decision
│ ↓
└── future runs ← merge ← PR ← approved change
RunTune needs permission to open a pull request, not to change what is deployed. Git remains the deployment boundary. Channels and deployment →
| from | artifact | what it writes |
|---|---|---|
| repeated failure | constraint | a guard rule, capped at what its failure rate supports (nudge, deny, or block) |
| repeated success | capability | a SKILL.md naming the commands that have worked, for code agents keep re-deriving |
| isolated, repeated work | sub-agent | a purpose-built agent type, granted exactly the tools its predecessors used |
| approved vs. ran vs. billed | route | a routing check or a retired clearance; a cheaper model always requires an eval first |
Each proposal is compiled to the narrowest artifact that works: a rule if a rule will do, a sub-agent only if a skill won't. Each proposal also records why the narrower options were rejected.
- There is one write boundary. The promoter is the only RunTune component that changes the
harness. Derivers, measurement and notifications can propose changes but have no code path
to apply them; a test enforces that only the two approval paths (the
applycommand and a parsed human reply) can reach the promoter. This is an architectural separation, not an instruction to a model. - The learner doesn't need the agent's authority. It only reads evidence: transcripts, the provider's bill, and a database it opens in read-only sessions (give it a SELECT-only role). Every change goes through a separate approval and promotion path.
- It can't widen its own boundaries. Loosening a rule, broadening a tool grant, or admitting a model needs a written reason, and for models a passing eval. None of these can be approved with a one-word reply.
- Its evidence sets a ceiling. A failure that happens 30% of the time can earn a nudge, never a block.
- It only edits what it wrote. It never overwrites a file a person wrote, and it refuses a proposal whose target changed after the proposal was reviewed.
- It stays out of host settings. The promoter refuses host settings, hook wiring and git internals as targets, whoever approves. Every decision goes into a hash-chained ledger.
- Ambiguous answers get a question back. "All but number three" is never read as "all".
RunTune was developed against about 238,000 recorded tool calls and model requests from one team's production agents: 150 days of Claude Code, Codex and OpenRouter traffic, plus 17,410 sub-agent runs. It found:
- repeated work that should be a reusable capability, including an existing CLI that served only 28% of the need it was built for;
- generic sub-agents doing work better suited to narrow, purpose-built ones;
- model traffic bypassing the approved router, which was traced to a helper and fixed;
- guard rules that helped, rules that stopped helping, and one that made failures worse until it was rewritten.
RunTune was wrong too. Five of its recommendations changed once checked against production:
- two apparent routing violations were a clearance-day transition and a benchmark;
- one "harmful" rule had already been fixed by its rewrite;
- a generated rule could never have fired;
- a generated skill named a command path that didn't exist.
Each mistake became a regression test, and the evidence doc keeps the wrong versions. A live experiment on one proposal came back inconclusive, and that result is published too.
Read the evidence and methodology →
Record → derive → gate → apply → review. Evidence comes from RunTune's own hook, a transcript backfill, the router's call log and the provider's bill. Every cluster is graded against its denominator, and anything withheld is reported with its reason. The promoter carries out only what a person approved, and review measures each change against a control.
Details: how it works · channels and deployment · warehouse sources · extending it.
Not a sandbox. Constraints run as tool-call hooks; for isolation, run agents inside an OS
sandbox, and RunTune inside that. The measurement caveats (observational rates, ok is not
quality, derivers are counters) are in the evidence doc.
RunTune merges three projects that each covered part of this loop. See docs/LINEAGE.md for what came from each, and why some ideas were changed or left out.
- CallusGuard: the failure side. Contributes failure tiers with action ceilings, the rule lifecycle, the probe-vs-prune distinction, and the hash-chained audit log, all ported. RunTune has its own recorder and enforcement hook and does not depend on CallusGuard. Its rulesets use the same shape, so either tool can enforce them.
- autoharness (MIT): the success side for Claude Code. Contributes the split where a model proposes and deterministic code writes, content-addressed evidence, use-rate against opportunity, merge versus death in the ledger, and the rule that RunTune only ever touches files it wrote.
- AutoRefine (arXiv 2601.22758): typed artifacts. Contributes the rule → skill → subagent compile order, correction plus preservation evidence, the replay gate, revisions that inherit their predecessors' cases, declared lineage, and suppressed gates that still record their verdict. These are reimplemented; no code was copied, because that repo ships without a license file.
See CONTRIBUTING.md. Results from your own agents, null results included, are welcome.
Apache-2.0. Copyright (C) 2026 Kai Karlstrom. See NOTICE for third-party credits.