Skip to content

Repository files navigation

ollaya

Run open decision models locally, the way Ollama runs LLMs.

Website · Models · Docs · Releases · Hugging Face

A decision model reads a state (a message, an email, a ticket, any JSON) plus typed questions (choice, score, noul) and returns calibrated probabilities in a single forward pass, in milliseconds. It never generates text. Ollaya pulls these models by name, serves them from a local daemon, and speaks TypeSafe's /v1/systemone wire format, so existing Jev clients work by changing one environment variable.

curl -fsSL https://ollaya.dev/install.sh | sh
ollaya run winnow:e4b --preset triage "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
intent            refund                                ███████████████░ 0.91
is_urgent         yes                                   ███████████████░ 0.92
frustration       2.89 / 3  very angry or using stron…  ██████████████░░ 0.86
refund_requested  yes                                   ████████████████ 0.99
churn_risk        yes                                   ████████████████ 0.99

winnow:e4b is the recommended model: 0.722 accuracy on typed decisions (TypeSafe's Jev: 0.738) and 89 ms for these five questions on an RTX 4090. It is a 4B-class language model, so without an NVIDIA GPU start with laya, which answers in a fraction of a second on a CPU. All models and their numbers: ollaya.dev/search.

Features

  • One binary. ollaya serve runs the daemon; ollaya run, pull, list, ps, show, rm, cp, stop and create work the way they do in Ollama. If the daemon isn't running, the CLI starts it.
  • TypeSafe-compatible. POST /v1/systemone, /v1/decisions and GET /v1/models are wire-identical to TypeSafe. The official SDK works unchanged when you set TYPESAFE_BASE_URL=http://localhost:11435.
  • Native API. /api/decide adds routing information and timings. /api/pull streams NDJSON progress, and there are /api/tags, /api/show, /api/ps and more. See docs/api.md.
  • Weights come from their authors. Ollaya publishes only small ONNX graphs, about 3 MB each. These graphs read the original weight files (usually model.safetensors) from the author's Hugging Face repository, pinned to a commit and verified by sha256. Models whose authors publish GGUF files (winnow, jevk5) run that file itself on llama.cpp. Ollaya never re-hosts weights.
  • For agents. ollaya mcp serves the models to Claude Code, Claude Desktop, Cursor and other MCP clients (claude mcp add ollaya -- ollaya mcp), and the ollaya-decisions skill teaches agents when and how to use them (npx skills add ollaya-dev/ollaya --skill ollaya-decisions).
  • Routers. laya detects the script and language of each request, then answers with laya:en or laya:multilingual.
  • Modelfiles. You can bake a question set into your own model:
    FROM laya
    QUESTIONS ./triage.json
    PARAMETER precision fp32
    
    Then run ollaya create triage -f Modelfile and ollaya run triage "…".
  • Fast and exact.
    • Hardware: ONNX Runtime on CPU, and CUDA on NVIDIA GPUs. GGUF models run on llama.cpp: CPU, CUDA, and Metal on Apple silicon.
    • Precision: fp16 on GPU and fp32 on CPU, chosen when the model loads.
    • Accuracy: fp32 exports give the same decision as the PyTorch reference on 100% of 2,383 questions per checkpoint.

Models

Model What it is
winnow:e4b Recommended. EldanRing's Winnow-E4B, a Gemma 4 fine-tune run from the author's Q8_0 GGUF on llama.cpp: 0.722 on typed-decisions (Jev: 0.738), 89 ms for five questions on an RTX 4090
laya Router: picks laya:en or laya:multilingual by language
laya:en English decision model (ModernBERT-large, 421M). The fastest: 8–10 ms for five questions on an RTX 4090
laya:multilingual 100+ languages (mmBERT-base, 322M)
laya:typed-decisions Fine-tuned on the typed-decisions workflows
decider, decider:4b, decider:0.8b Mapika's Qwen3.5 decoders, 2B (the default), 4B and 0.8B: 0.680 on typed-decisions for 4B, 0.591 for 2B
decider:2b-vision Mapika's Qwen3.5-2B vision-language decider: questions about an image (--image, images on /api/decide) as well as the state
kev, kev:0.8b, kev:9b Jared Palmer's Kev: a LoRA and a pointer head on Qwen3.5 (4B by default, 0.8B, 9B), calibrated. kev:4b scores 0.669 on typed-decisions and kev:9b 0.722, as much as winnow:e4b
decision Decision 1.0 Eos by the vLLM Semantic Router contributors: a fine-tuned Qwen3.5-0.8B with an endpoint head, 16k-token rows
qwen3guard Qwen3Guard-Gen-0.6B safety guard with built-in questions: safe, controversial or unsafe, and the category
nli, nli:modernbert-large Moritz Laurer's zero-shot NLI classifiers (DeBERTa-v3-large, ModernBERT-large)
gliclass Knowledgator's instruction-following zero-shot classifier (DeBERTa-v3-large)
von Victor Hugo Panisa's Von 1.1 (ModernBERT-large): every option scored at its own marker, 8k-token context
winnow EldanRing's Winnow-12B, the larger sibling of winnow:e4b: 0.702 on typed-decisions
clm Contrastive-LM's CLM-v0.1-8B: the Qwen3-8B encoder and two projection heads score options by similarity, with questions and options cached. 0.357 on typed-decisions; built for agent, game and tool-calling states
jevk5 alibiserikbay's JevK5 v0.3, a Qwen3.5-4B fine-tune run from the author's Q8_0 GGUF on llama.cpp, up to 16 options

Browse them at ollaya.dev/search. Laya tags ending in -fp32 or -fp16 pin the precision. The derived files of every model are also published at huggingface.co/ollaya-dev.

Install

  • Linux (x86_64 or arm64, glibc ≥ 2.38, e.g. Ubuntu 24.04+): curl -fsSL https://ollaya.dev/install.sh | sh. When an NVIDIA GPU is present (driver R525+), the installer adds the CUDA runtime: CUDA 13 for R580+, CUDA 12 for older drivers.
  • macOS (Apple silicon): the same command.
  • Windows (x64): irm https://ollaya.dev/install.ps1 | iex in PowerShell. When an NVIDIA GPU is present (driver R527+), the installer adds the CUDA runtime, as on Linux.
  • Desktop app for macOS, Windows and Linux: start and stop the server, download models and run them in one window. On macOS it lives in the menu bar. Get it from ollaya.dev/download.
  • Docker: docker run -d --gpus=all -p 11435:11435 ghcr.io/ollaya-dev/ollaya:cuda (:cuda12 for host drivers older than R580), or ghcr.io/ollaya-dev/ollaya for CPU only.

Configuration is through environment variables: OLLAYA_HOST, OLLAYA_MODELS, OLLAYA_KEEP_ALIVE, OLLAYA_DEVICE, OLLAYA_API_KEY and others, listed in docs/api.md §15.

Repository

Path What
crates/ollaya The binary: CLI, daemon, runner
crates/ollaya-server HTTP API, scheduler (one runner process per model), model resolution
crates/ollaya-api API types and client; the contract is docs/api.md
crates/ollaya-registry Model names, manifests, blob store, resumable pulls
crates/ollaya-decision Question schema, sequence layouts, calibration, answers
crates/ollaya-runner Inference engines (ONNX Runtime, and llama.cpp for GGUF models)
crates/ollaya-lang Script and language detection for routers
convert/ Build-time Python: ONNX export, parity checks, packaging
site/ The website and the static model registry host

Development

cargo test --workspace
cargo build --release -p ollaya --features cuda   # CUDA build (x86-64 Linux and Windows)

convert/ rebuilds models. It exports them, checks parity against the PyTorch reference, generates golden fixtures, and packages the result into registry/. See the module docstrings. cd convert && uv sync installs it: with CUDA 13 torch on Linux and Windows, and with the CPU and MPS build from PyPI on Apple silicon Macs, where exports and parity run on the CPU.

License

Apache-2.0. Each model keeps its own license: laya (Convai Innovations), decider (Mapika), kev (Jared Palmer, on Qwen3.5 by the Qwen team), decision (the vLLM Semantic Router contributors, on Qwen3.5), qwen3guard (Qwen team), gliclass (Knowledgator), von (Victor Hugo Panisa), winnow (EldanRing, on Gemma 4 by Google DeepMind), jevk5 (alibiserikbay, on Qwen3.5) and nli:modernbert-large are Apache-2.0, and nli:deberta-v3-large (Moritz Laurer) is MIT. llama.cpp, which Ollaya ships for GGUF models, is MIT.

Ollaya is an independent project. It is not affiliated with or endorsed by Ollama or TypeSafe.

About

Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.

Topics

Resources

Stars

998 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages