A dependency-free Node.js runtime that lets a human hand off a high-level software task ("fix X in project Y") to a chain of specialized Claude Code agents — Orchestrator → Engineer → QA → Reviewer → Context Curator — and get back a verified result, without babysitting each handoff by hand.
The interesting part isn't the happy path. It's what happens when an agent lies about what it did, when a repo is dirty, when a run gets interrupted mid-flight, or when a retry loop should have stopped three cycles ago. This project treats those as the default case to design for, not edge cases to patch later.
Runtime is real, tested (94 automated tests, node --test), and has been
exercised end-to-end against a real codebase across every failure mode below
— dirty-worktree blocking, forced QA-fail retry, forced review-rejection
retry, interruption/resume, and a real gated local commit, each independently
verified against actual git state, not agent self-reports.
Not yet claimed: a 10–20 run production-stability soak (the project's own roadmap calls this "Phase 13") is intentionally still pending — this repo does not overstate validation it hasn't done. What's here is honest: a solid, tested core, proven under adversarial conditions, not yet proven at volume.
Handing a coding agent a task and trusting its own account of what happened doesn't scale past toy examples. Three failure modes showed up immediately in real use:
- Agents misreport their own work. In real runs, an agent twice claimed it wrote a file that didn't exist. Not malice — just unreliable self-narration, the same way a person might misremember what they just did.
- A shared working tree has state that isn't yours to touch. An automated agent that can't tell "my change" from "the user's uncommitted work from an hour ago" will eventually stage or overwrite the wrong thing.
- Retry loops need a hard ceiling, not a suggestion. Without an enforced limit, a failing fix attempt can loop indefinitely or, worse, an agent can route around the guardrail entirely by calling a lower-level primitive directly.
Each of these has a concrete, tested countermeasure below.
request
│
▼
Orchestrator ──resolves project──▶ project_registry.mjs
│ classifies task, creates a run, checks dirty-worktree baseline
▼
Engineer ──implements──▶ QA ──verifies independently──▶ [retry ≤3x on fail]
│ │
│ pass ▼
│ Reviewer (COMPLEX/CRITICAL only) ──▶ [retry ≤2x on changes_required]
│ │ approved
▼ ▼
Context Curator ──proposes context update──▶ zero-trust verification gate ──▶ archive ──▶ report
| Agent | Job | Cannot do |
|---|---|---|
| Orchestrator | Resolve project, classify task, create/track the run, dispatch the next agent, enforce retry ceilings | Write to the knowledge base, commit, push, merge, deploy |
| Engineer | Audit → root cause → minimal change → local test → report, inside the target repo | Decide it's done — QA verifies independently |
| QA | Independently re-verify the Engineer's work against acceptance criteria | Fix code itself |
| Reviewer | Correctness, scope creep, security, architecture — only for COMPLEX/CRITICAL tasks | Approve its own fixes in a loop without a ceiling |
| Context Curator | Propose what durable knowledge changed; the only agent that can trigger a local commit | Push, merge, deploy, touch anything outside an explicit allowlist |
Every run is a JSON file (state.json) plus an append-only event log
(events.jsonl), always in a small set of explicit states
(CREATED → IMPLEMENTING → TESTING → [FIXING → TESTING]* → REVIEWING → [FIXING → ...]* → CONTEXT_UPDATE → COMPLETED,
or BLOCKED/FAILED/CANCELLED/NEEDS_APPROVAL). retry_policy.mjs
enforces a hard ceiling on both loops — exceed it and the run stops in
NEEDS_APPROVAL with a structured report, never a silent infinite loop.
If the process is interrupted mid-run, resume.mjs reconstructs exactly
what already happened from the persisted state and event log — not from
memory, not from a prompt re-describing what it thinks happened — and
returns the one safe next step. A CLI guard sits in front of every
agent-dispatch path and refuses to re-run a step the resume plan says is
already done, so re-invoking a finished step and duplicating its side effect
isn't just discouraged, it's blocked in code.
Before touching a shared repo, dirty_worktree.mjs captures what's already
uncommitted. If the run's target overlaps with pre-existing uncommitted
work, the run stops at NEEDS_APPROVAL — it never resets, stashes, or
checks out over a human's in-progress changes.
This is the part most agent frameworks skip. An agent's claim about what it
changed is a proposal, never a fact. claim_verification.mjs
independently re-derives verified_actual from real git/filesystem state
and diffs it against agent_claimed. A relevant mismatch — a claimed file
that doesn't exist, a claimed "no changes" when a real diff exists, a
learning claimed as written but nowhere on disk — logs REPORT_MISMATCH and
fails closed to NEEDS_APPROVAL, before anything is committed.
Only the Context Curator can trigger a commit, and only through one exported
function (proposeAndCommit) that enforces the full sequence in code —
not by trusting the calling agent to follow instructions:
baseline capture → verified-diff (git, not the claim) → mismatch check
→ secret scan (on real file content) → allowlist check
→ selective `git add <path>` (never `-A`/`.`) → staged-diff verification
→ commit (local only — never push, never merge)
Every lower-level step is a private function; there is no shortcut path that
reaches git commit without going through all of them. Dry-run by default.
Every event written to events.jsonl, every learning candidate, and every
file considered for commit is scanned for secret-shaped strings
(validators.mjs) and redacted before it ever touches disk.
agents/ Canonical agent definitions (Markdown + YAML frontmatter)
runtime/ Zero-dependency Node.js/ESM runtime + full test suite
workflows/ Per-task-type routing (bug-fix / implementation / audit)
evals/ 8 scenario definitions with expected states/agents/side-effects
decisions/ Architectural decisions made building this, with reasoning
learnings/ Append-only, never auto-promoted, agent-facing pattern notes
examples/ Placeholder target projects + an example project registry
docs/ Design spec this runtime implements
cd runtime
npm test # 94 tests, node's built-in test runner, zero dependenciesTo point the Orchestrator at real projects, edit projects/registry.json
(alias → path → local instructions) and drop a copy of agents/*.md into
your Claude Code agents directory.
See decisions/global/ — in particular
0002, the
concrete incident that motivated the zero-trust verification layer.
Agent definitions in agents/ and some inline code comments are
in Italian — the language they were authored and battle-tested in. The
runtime code, tests, and this README are in English. Behavior and test
coverage are unaffected either way; nothing important is locked behind the
Italian text.
MIT.