Run open decision models locally, the way Ollama runs LLMs.
Website · Models · Docs · Releases · Hugging Face
A decision model reads a state (a message, an email, a ticket, any JSON) plus typed questions
(choice, score, noul) and returns calibrated probabilities in a single forward pass, in
milliseconds. It never generates text. Ollaya pulls these models by name, serves them from a
local daemon, and speaks TypeSafe's /v1/systemone wire format, so existing Jev clients work by
changing one environment variable.
curl -fsSL https://ollaya.dev/install.sh | sh
ollaya run winnow:e4b --preset triage "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."intent refund ███████████████░ 0.91
is_urgent yes ███████████████░ 0.92
frustration 2.89 / 3 very angry or using stron… ██████████████░░ 0.86
refund_requested yes ████████████████ 0.99
churn_risk yes ████████████████ 0.99
winnow:e4b is the recommended model: 0.722 accuracy on typed decisions (TypeSafe's Jev: 0.738)
and 89 ms for these five questions on an RTX 4090. It is a 4B-class language model, so without an
NVIDIA GPU start with laya, which answers in a fraction of a second on a CPU. All models and their
numbers: ollaya.dev/search.
- One binary.
ollaya serveruns the daemon;ollaya run,pull,list,ps,show,rm,cp,stopandcreatework the way they do in Ollama. If the daemon isn't running, the CLI starts it. - TypeSafe-compatible.
POST /v1/systemone,/v1/decisionsandGET /v1/modelsare wire-identical to TypeSafe. The official SDK works unchanged when you setTYPESAFE_BASE_URL=http://localhost:11435. - Native API.
/api/decideadds routing information and timings./api/pullstreams NDJSON progress, and there are/api/tags,/api/show,/api/psand more. See docs/api.md. - Weights come from their authors. Ollaya publishes only small ONNX graphs, about 3 MB each.
These graphs read the original weight files (usually
model.safetensors) from the author's Hugging Face repository, pinned to a commit and verified by sha256. Models whose authors publish GGUF files (winnow,jevk5) run that file itself on llama.cpp. Ollaya never re-hosts weights. - For agents.
ollaya mcpserves the models to Claude Code, Claude Desktop, Cursor and other MCP clients (claude mcp add ollaya -- ollaya mcp), and theollaya-decisionsskill teaches agents when and how to use them (npx skills add ollaya-dev/ollaya --skill ollaya-decisions). - Routers.
layadetects the script and language of each request, then answers withlaya:enorlaya:multilingual. - Modelfiles. You can bake a question set into your own model:
Then run
FROM laya QUESTIONS ./triage.json PARAMETER precision fp32ollaya create triage -f Modelfileandollaya run triage "…". - Fast and exact.
- Hardware: ONNX Runtime on CPU, and CUDA on NVIDIA GPUs. GGUF models run on llama.cpp: CPU, CUDA, and Metal on Apple silicon.
- Precision: fp16 on GPU and fp32 on CPU, chosen when the model loads.
- Accuracy: fp32 exports give the same decision as the PyTorch reference on 100% of 2,383 questions per checkpoint.
| Model | What it is |
|---|---|
winnow:e4b |
Recommended. EldanRing's Winnow-E4B, a Gemma 4 fine-tune run from the author's Q8_0 GGUF on llama.cpp: 0.722 on typed-decisions (Jev: 0.738), 89 ms for five questions on an RTX 4090 |
laya |
Router: picks laya:en or laya:multilingual by language |
laya:en |
English decision model (ModernBERT-large, 421M). The fastest: 8–10 ms for five questions on an RTX 4090 |
laya:multilingual |
100+ languages (mmBERT-base, 322M) |
laya:typed-decisions |
Fine-tuned on the typed-decisions workflows |
decider, decider:4b, decider:0.8b |
Mapika's Qwen3.5 decoders, 2B (the default), 4B and 0.8B: 0.680 on typed-decisions for 4B, 0.591 for 2B |
decider:2b-vision |
Mapika's Qwen3.5-2B vision-language decider: questions about an image (--image, images on /api/decide) as well as the state |
kev, kev:0.8b, kev:9b |
Jared Palmer's Kev: a LoRA and a pointer head on Qwen3.5 (4B by default, 0.8B, 9B), calibrated. kev:4b scores 0.669 on typed-decisions and kev:9b 0.722, as much as winnow:e4b |
decision |
Decision 1.0 Eos by the vLLM Semantic Router contributors: a fine-tuned Qwen3.5-0.8B with an endpoint head, 16k-token rows |
qwen3guard |
Qwen3Guard-Gen-0.6B safety guard with built-in questions: safe, controversial or unsafe, and the category |
nli, nli:modernbert-large |
Moritz Laurer's zero-shot NLI classifiers (DeBERTa-v3-large, ModernBERT-large) |
gliclass |
Knowledgator's instruction-following zero-shot classifier (DeBERTa-v3-large) |
von |
Victor Hugo Panisa's Von 1.1 (ModernBERT-large): every option scored at its own marker, 8k-token context |
winnow |
EldanRing's Winnow-12B, the larger sibling of winnow:e4b: 0.702 on typed-decisions |
clm |
Contrastive-LM's CLM-v0.1-8B: the Qwen3-8B encoder and two projection heads score options by similarity, with questions and options cached. 0.357 on typed-decisions; built for agent, game and tool-calling states |
jevk5 |
alibiserikbay's JevK5 v0.3, a Qwen3.5-4B fine-tune run from the author's Q8_0 GGUF on llama.cpp, up to 16 options |
Browse them at ollaya.dev/search. Laya tags ending in
-fp32 or -fp16 pin the precision. The derived files of every model are also published at
huggingface.co/ollaya-dev.
- Linux (x86_64 or arm64, glibc ≥ 2.38, e.g. Ubuntu 24.04+):
curl -fsSL https://ollaya.dev/install.sh | sh. When an NVIDIA GPU is present (driver R525+), the installer adds the CUDA runtime: CUDA 13 for R580+, CUDA 12 for older drivers. - macOS (Apple silicon): the same command.
- Windows (x64):
irm https://ollaya.dev/install.ps1 | iexin PowerShell. When an NVIDIA GPU is present (driver R527+), the installer adds the CUDA runtime, as on Linux. - Desktop app for macOS, Windows and Linux: start and stop the server, download models and run them in one window. On macOS it lives in the menu bar. Get it from ollaya.dev/download.
- Docker:
docker run -d --gpus=all -p 11435:11435 ghcr.io/ollaya-dev/ollaya:cuda(:cuda12for host drivers older than R580), orghcr.io/ollaya-dev/ollayafor CPU only.
Configuration is through environment variables: OLLAYA_HOST, OLLAYA_MODELS,
OLLAYA_KEEP_ALIVE, OLLAYA_DEVICE, OLLAYA_API_KEY and others, listed in
docs/api.md §15.
| Path | What |
|---|---|
crates/ollaya |
The binary: CLI, daemon, runner |
crates/ollaya-server |
HTTP API, scheduler (one runner process per model), model resolution |
crates/ollaya-api |
API types and client; the contract is docs/api.md |
crates/ollaya-registry |
Model names, manifests, blob store, resumable pulls |
crates/ollaya-decision |
Question schema, sequence layouts, calibration, answers |
crates/ollaya-runner |
Inference engines (ONNX Runtime, and llama.cpp for GGUF models) |
crates/ollaya-lang |
Script and language detection for routers |
convert/ |
Build-time Python: ONNX export, parity checks, packaging |
site/ |
The website and the static model registry host |
cargo test --workspace
cargo build --release -p ollaya --features cuda # CUDA build (x86-64 Linux and Windows)convert/ rebuilds models. It exports them, checks parity against the PyTorch reference,
generates golden fixtures, and packages the result into registry/. See the module docstrings.
cd convert && uv sync installs it: with CUDA 13 torch on Linux and Windows, and with the CPU and
MPS build from PyPI on Apple silicon Macs, where exports and parity run on the CPU.
Apache-2.0. Each model keeps its own license: laya (Convai Innovations), decider (Mapika),
kev (Jared Palmer, on Qwen3.5 by the Qwen team), decision (the vLLM Semantic Router
contributors, on Qwen3.5), qwen3guard (Qwen team), gliclass (Knowledgator), von (Victor Hugo
Panisa), winnow (EldanRing, on Gemma 4 by Google DeepMind), jevk5 (alibiserikbay, on Qwen3.5)
and nli:modernbert-large are Apache-2.0, and nli:deberta-v3-large (Moritz Laurer) is MIT. llama.cpp, which Ollaya ships for
GGUF models, is MIT.
Ollaya is an independent project. It is not affiliated with or endorsed by Ollama or TypeSafe.