A local decision layer for agents and workflows. You hand it a state (text or JSON) and a schema of named questions, each with a fixed list of allowed answers. It hands back a typed answer per question, a confidence score, which model decided, and how long it took — running entirely on your machine, on two local models:
- Fast layer: MiniCPM5-1B (Apache-2.0) answers every question in one call, thinking mode off.
- Deep layer: Spark-X2.5-4B (Apache-2.0) re-answers only the questions the fast layer wasn't confident about.
Both run on llama-server from llama.cpp (MIT). Spark-X2.5
needs the Spark2_5ForCausalLM architecture, added to the XHToken/llama.cpp
fork first and since merged into mainline llama.cpp
(ggml-org/llama.cpp#27868, released from build
b10828 on). There is no cloud dependency and no telemetry: the only network calls reflex makes are
to Hugging Face (model weights) and GitHub (the llama.cpp runtime), both of which you can point at
your own mirrors via HF_ENDPOINT.
reflex is a working title, kept in one constant (TOOL_NAME in src/constants.ts) so it's a
one-line change to rename.
- Linux, macOS, or Windows.
- Linux/macOS: builds the XHToken fork from source by default (needs git, CMake, and a C++
compiler —
reflex setupchecks for these and prints install instructions if anything's missing), falling back to a downloaded prebuilt binary if no compiler is found. - Windows: always downloads a prebuilt
llama-serverfrom the official ggml-org/llama.cpp releases instead of trying to automate an MSVC/clang build. This works because mainline llama.cpp now has native Spark2_5 support (see above) — no fork-specific binary is needed.reflexalways resolves the latest compatible release rather than a pinned one.
- Linux/macOS: builds the XHToken fork from source by default (needs git, CMake, and a C++
compiler —
- Memory:
reflex upestimates each model's RAM need asfile size × 1.5(a heuristic, not a benchmark) and warns before starting both models if that clearly won't fit, offering to start the fast layer alone or use the smallerspark-x2.5-1.7bdeep model instead. - Building from source: an NVIDIA GPU with
nvcconPATHis used automatically (-DGGML_CUDA=ON); Apple Silicon gets Metal by default; otherwise it's CPU-only. The Windows/ fallback prebuilt path is CPU-only — pointruntime.fast.binary/runtime.deep.binaryat a CUDA/Vulkan build from the same releases page yourself if you want GPU acceleration there.
pnpm install
pnpm build
node dist/cli/index.js setup # or: npm link, then `reflex setup`setup checks prerequisites, clones and builds the llama.cpp fork, downloads both models (after
showing you their size and license — pass --yes to skip the prompt), and finishes by running
doctor.
reflex up # start both model servers
reflex decide --schema examples/schema.json --state-file examples/state.txt --pretty
reflex down # stop themdecide autostarts the servers if autostart is enabled in the config (the default), so up is
optional for a one-off call. Real output from this exact example (Windows, RTX 4090, CUDA build):
QUESTION ANSWER CONFIDENCE BY LATENCY
sentiment negative 0.807 fast 516ms
priority high 1.000 deep 208ms
needs_human yes 1.000 deep 309ms
Total: 859ms, escalation rate: 67%
State can also be JSON (examples/state.json) or piped via stdin. --pretty prints a table;
without it, decide prints the same result as JSON, suitable for piping into another program.
| Command | What it does |
|---|---|
reflex setup |
prerequisites → build runtime → pull both models → doctor |
reflex models list | pull <name> | rm <name> |
manage local model weights |
reflex up [--fast-only|--deep-only] [--force] |
start model server(s) |
reflex down |
stop model server(s) |
reflex status [--json] |
show ports, pids, health |
reflex doctor |
prerequisites, binary, models, server health, structured-output/logprobs/ping checks |
reflex decide --schema <file> [--state <text>|--state-file <file>] [--pretty] |
one decision |
reflex serve [--port <n>] |
POST /decide, GET /health, plus a GET / test page, bound to localhost only |
reflex bench --tasks <file> [--pretty] |
accuracy/p50/p95/escalation-rate/ECE, fast-only vs. deep-only vs. combined |
reflex calibrate --tasks <file> |
fits router.calibration.temperature against labeled tasks, saves it to config |
A tasks.jsonl file (see examples/tasks.jsonl) has one JSON object per line:
{"schema": {...}, "state": ..., "expected": {"questionName": "answer"}}.
All fields are optional; shown values are the defaults. runtime.fast.binary / runtime.deep.binary
have no default — omit them entirely unless you want to override the built binary (see below).
{
"models": { "fast": "minicpm5-1b", "deep": "spark-x2.5-4b" },
"runtime": {
"contextSize": { "fast": 4096, "deep": 8192 },
"gpuLayers": "auto",
"threads": -1
},
"fast": { "temperature": 0.7, "topP": 0.95, "topK": 40, "minP": 0.0 },
"deep": { "temperature": 1.0, "topP": 0.95, "topK": -1, "minP": 0.0 },
"router": {
"threshold": 0.8,
"onMissingConfidence": "escalate",
"calibration": { "temperature": 1.0 }
},
"ports": { "fast": 8081, "deep": 8082 },
"server": { "host": "127.0.0.1", "port": 8787 },
"autostart": true
}runtime.contextSize.deepis kept small on purpose — Spark-X2.5's native 1M-token context needs a lot of memory; raise it if you actually need long context and have the RAM for it.runtime.fast.binarylets you point the fast layer at a differentllama-serverbuild (e.g. a prebuilt one from ggml-org) if the XHToken fork ever doesn't run MiniCPM5 cleanly;doctorchecks for this.runtime.deep.binaryexists for symmetry but the deep layer normally needs the fork'sSpark2_5ForCausalLMsupport.fast/deepgeneration defaults come straight from each model's card (MiniCPM5's no-think recommendation; Spark-X2.5's README), includingminP: 0.0— the MiniCPM llama.cpp cookbook notes the library default ofmin_p=0.05can suppress the exact tokens needed to break a repetition loop.router.onMissingConfidence: if a backend doesn't return usable logprobs, confidence isnull;"escalate"(default) sends that question to the deep layer anyway,"accept"trusts the fast answer as-is.- Respects
HF_TOKEN(gated/private repos) andHF_ENDPOINT(mirrors) from the environment, same as the Hugging Face CLI.
Only the two registry models (plus spark-x2.5-1.7b, offered as a lower-memory deep-layer
fallback) can be pulled by name. reflex models pull <name> --unsafe-repo <owner/repo> bypasses
the registry for anything else, at your own risk (license/architecture aren't verified).
- The fast layer answers every question in one call, with JSON-schema-constrained decoding
(
json_schemaon llama-server's native/completionendpoint) and thinking mode off. - Per question, reflex finds the answer's token span in the generated JSON and computes the joint
probability of those tokens, temperature-scaled by
router.calibration.temperature(reflex calibratefits this against labeled tasks). - Below
router.threshold, or if calibration says the answer isn't a good match, that question alone is escalated to the deep layer (thinking mode on) — not the whole schema.
Read this before trusting a number reflex prints.
-
The only numbers below are from real
reflex benchruns against the real models — onexamples/tasks.jsonl(6 tasks, 18 questions total), Windows, RTX 4090, CUDA build, both models at their config defaults. Two consecutive runs:mode accuracy p50 p95 escalation ECE fast-only (run 1) 0.39 126ms 355ms 0% 0.32 fast-only (run 2) 0.50 106ms 307ms 0% 0.14 deep-only (run 1) 0.78 361ms 365ms 0% 0.20 deep-only (run 2) 0.78 370ms 390ms 0% 0.20 combined (run 1) 0.67 493ms 519ms 83% 0.29 combined (run 2) 0.72 533ms 547ms 89% 0.25 Three things this small run actually shows, and doesn't:
- Accuracy varies noticeably run to run (fast-only: 0.39 vs 0.50) because both layers sample at non-zero temperature (0.7 fast, 1.0 deep, per each model's own card) and 18 questions is a tiny sample — one flipped answer moves accuracy by ~5.6 points. Don't read a single run's number as precise.
combinedscored belowdeep-onlyalone on this task set, even while escalating 83-89% of questions. That's not a bug: it means the uncalibrated 0.8 threshold let a few genuinely wrong fast answers through as "confident enough." This is exactly whatreflex calibrateexists to fix — it wasn't run here, sorouter.calibration.temperatureis still the un-fit default 1.0.- This says nothing about accuracy on your own task distribution, or about the models' quality in
general — it's 6 illustrative support-ticket examples on one machine. Run
reflex bench --tasks <your own tasks.jsonl> --pretty(andreflex calibratebefore trusting the escalation threshold) against your own data before drawing conclusions.
-
Confidence is an approximation, not a full-vocabulary probability. llama-server's
n_probsonly returns the top-N alternatives it sampled from at each position, not the full vocabulary distribution. Temperature scaling here renormalizes and rescales within that top-N set, per token, then takes the product across the tokens spanning the answer. This is the honest thing to do with the data actually available, but it is not equivalent to textbook temperature scaling over the full softmax. -
n_probson/completionis confirmed by the fork's own docs; on/v1/chat/completionsit is not documented at all. That's why reflex builds the prompt via/apply-templateand calls the native/completionendpoint instead of the OpenAI-compatible chat endpoint. If a future backend swap doesn't honorn_probs, confidence extraction returnsnullfor that answer — this is treated as a real, expected outcome (router.onMissingConfidence), not a bug to paper over. -
Confidence is fast-layer-only for routing and calibration. The deep layer's confidence is computed and reported the same way, but nothing thresholds on it — it never escalates further.
reflex calibratealso only calibrates against fast-layer confidence, since that's whatrouter.thresholdactually uses. -
Memory estimates are a heuristic (
file size × 1.5), not measured. KV cache size depends on context length and how many slots llama-server allocates; if you setruntime.contextSize.deepmuch higher than the 8192 default, the real requirement will be higher than this estimate says. -
The prebuilt/fallback path always resolves the latest ggml-org/llama.cpp release, not a pinned version, since these nightly-style
bNNNNbuilds are meant to always work standalone. This means the exact binary in use can change between runs on a machine that uses the fallback path (recorded each time inlock.jsonfor traceability) — a deliberate tradeoff for not having to maintain a version pin, per how these releases are meant to be consumed. -
Tests never load a real model — they run against small local HTTP servers standing in for
llama-server's documented endpoint shapes (/apply-template,/completionwithcompletion_probabilities), a fake Hugging Face API for model downloads, and a fake GitHub API for the prebuilt-binary path (verified for real, including actualtar-based zip/tar.gz extraction, in addition to the mocked tests).
pnpm install
pnpm typecheck
pnpm test
pnpm buildLayout: src/cli (thin command wiring), src/core (decide/router/confidence/calibration — the
importable library surface, import { decide } from "reflex"), src/models (registry/download/
cache), src/runtime (build the fork, spawn/stop llama-server, orchestration), src/bench,
src/server (the localhost HTTP API), tests, examples.
Apache-2.0. See LICENSE.