Local semantic search, browse, and resume over your Claude Code and Codex transcripts. Everything runs on-device — embeddings, reranking, and the vector index never leave your machine.
You talk to a lot of agents. santa makes that history searchable: "that postgres migration we argued about", "the auth refactor from last week" — find the session, read it, and resume it where you left off.
For the product's purpose, scope, and ownership boundaries, see the current vision. The documentation index separates current product documents from contextual references.
The main way to use santa is its full-screen TUI:
santa tui # browse + search your whole history, resume with one keyEverything it does is also a plain command, for scripting or muscle memory:
santa query "the flaky test we kept fighting" # one-shot hybrid search
santa recent # recent sessions, grouped by project
santa refresh # re-index (the TUI also does this on launch)Needs the .NET 10 SDK (the build publishes a single-file binary; the runtime is framework-dependent, so it's small).
./install.sh # GPU ONNX Runtime by default
santa models download
santa refreshsanta models download pulls two ONNX models (embedding + reranker) from HuggingFace
and the sqlite-vec extension from GitHub,
then works fully offline. CUDA is the default inference provider.
Before loading a CUDA model, Santa requires three consecutive idle readings from
nvidia-smi (at most 10% utilization and 25% VRAM in use). Santa and Shepherd share
~/.local/state/agent-tooling/gpu-inference.lock, so their inference passes cannot
compete. A busy or unavailable GPU defers the command; it never silently falls back
to CPU. Use an explicit CPU build and provider when wanted:
SANTA_ONNX_RUNTIME_FLAVOR=cpu ./install.sh
santa devices --provider cpu
santa refresh --provider cpuCUDA embedding is memory-bounded: batches are limited by their padded token area, the embedding allocator is capped at 7 GiB, and the reranker allocator at 2 GiB. This keeps a single maximum-length transcript turn from padding a large batch and exhausting a 12 GiB GPU.
If CUDA initialization fails, the command fails rather than silently continuing with
BM25. Use --provider keyword-only when BM25-only operation is what you intended.
An hourly cron keeps the index (and therefore session titles, which cockpit reads) current:
0 * * * * PATH=$HOME/.local/bin:/usr/bin:/bin:/usr/lib/wsl/lib santa refresh --quiet --provider cuda >> ~/.local/share/santa/refresh.log 2>&1
/usr/lib/wsl/lib is load-bearing on WSL: that is where nvidia-smi lives, and the CUDA
courtesy gate skips the entire run — summarisation included — when it cannot find it.
Omitting the path silently stops all indexing and every downstream title goes stale; check
refresh.log for skipped reason=nvidia-smi-unavailable if titles stop updating.
santa tui is the main interface — a full-screen, keyboard-driven browser over your
whole session history. It runs a silent incremental index on launch, so it's always
current; you rarely need santa refresh by hand.
Two tabs, switched with Tab:
- Browse — every session as a scrollable list (date · cwd · turns · duration ·
title), newest first. A green
●marks sessions running right now;✓marks ones you've completed;cxtags Codex sessions. A detail pane shows the summary. - Search — hybrid BM25 + on-device vector search with a reranker. Type a query,
hit Enter; the match panel shows the snippet and the
bm25/vec/fused/rerankscores.
| Key | Browse | Search |
|---|---|---|
↑ ↓ · PgUp PgDn · Home End |
move selection | move through results |
Tab |
switch tab | switch tab |
/ |
live filter (title · cwd · branch · summary) | edit the query |
Enter |
expand the detail pane | run the search |
s |
toggle the detail pane | — |
c |
mark session completed (hide from default search) | — |
r |
resume → cockpit if attached, else a new terminal | resume the hit |
t |
resume in a new terminal (ignore the cockpit handoff) | same |
v |
toggle stacked ↔ side-by-side layout (≥140 cols) | same |
Ctrl-O |
theme picker (live preview, Enter saves) |
same |
Ctrl-Q |
quit | quit |
Resume is the payoff: land on a session, press r, and you're back in it — a fresh
terminal tab, or dropped straight into a cockpit
pane if you launched the TUI from there (cockpit --santa). Theme and layout choices
persist across runs.
Flags: --provider cpu|cuda|keyword-only · --keyword-only (compatibility alias
for BM25-only) · --no-refresh (skip the launch-time index) · --device <id>
(pick a GPU when using CUDA).
Everything — the index DB, the downloaded models, and native libs — lives under
~/.local/share/santa/ (XDG-respecting). Override the whole tree with SANTA_HOME.
The index is built from your transcripts and is never committed or shared.
Upgrading from a
santa-claudeinstall?install.shmigrates the old~/.local/share/santa-claude/state dir automatically, and the binary still honoursSANTA_CLAUDE_HOMEas a deprecated alias forSANTA_HOME.
Summarisation and classification are routed by SANTA_SUMMARIZER, and each backend names
its model explicitly — neither path inherits the underlying CLI's own default.
| Variable | Default | Does |
|---|---|---|
SANTA_SUMMARIZER |
codex |
Backend for tool-less work (summaries, non-agent classification): codex or claude. Agent-mode recipes always run on Claude regardless. |
SANTA_CLAUDE_MODEL |
claude-haiku-4-5 |
Model for the claude -p path when a recipe or --model names none. A recipe that does name one still wins. |
SANTA_CODEX_MODEL |
gpt-5.5 |
Model for the codex exec path. Overrides the per-request model outright — Claude model names mean nothing to Codex. |
--model is always passed to claude -p. An empty value resolves to SANTA_CLAUDE_MODEL
rather than falling through to whatever default the CLI happens to ship, so the model in the
classifications and summaries tables is always the one that actually ran.
Everything the TUI does, minus the UI — for piping, scripts, or muscle memory.
| Command | Does |
|---|---|
santa tui |
The full-screen interface above. |
santa refresh |
Incrementally index new/changed sessions; CUDA inference when the GPU is idle. |
santa query <text> |
Hybrid semantic + keyword search across all history. |
santa recent |
Recently-touched sessions, grouped by project. |
santa related <id> |
Sessions semantically near a given one. |
santa show <id> / export <id> |
Read or dump a full transcript. |
santa resume <id> |
Re-open a session in its original CLI (Claude or Codex). |
santa stands alone, but it's built to compose with two siblings:
- cockpit — a tmux control surface that
resumes several sessions at once into one live grid.
santa resumecan hand a session straight to it. - agent-fusion — a multi-agent harness that runs Claude + Codex on one task and fuses the output.
See the agent-tooling umbrella for how the three fit together.
Personal tooling, shared as-is. No tests, no CI. Built and run on WSL + Linux.
The build surfaces an NU1903 advisory on a transitive SQLitePCLRaw native
package — noted, not yet bumped. PRs welcome; expectations modest.
The Unlicense — released into the public domain. Do whatever you want.