Run OpenClaw on a local model. Zero API cost. One OpenAI-compatible endpoint over Ollama, llama.cpp, vLLM or LM Studio — plus AirLLM for models bigger than your GPU.
-
Updated
Aug 31, 2026 - Python
Run OpenClaw on a local model. Zero API cost. One OpenAI-compatible endpoint over Ollama, llama.cpp, vLLM or LM Studio — plus AirLLM for models bigger than your GPU.
Run 70B+ LLMs on a single 4GB GPU — no quantization required.
OpenAI-compatible API server wrapping AirLLM (github.com/lyogavin/airllm) — run large LLMs on small GPUs, CPU, Apple silicon or Docker
Custom ML architectures, training pipelines, and Dockerized deployment experiments — built from scratch. Ongoing.
PyTorch inference acceleration engine: rejection-sampling speculative decoding (1.92x speedup), 5-way quantization bake-off & AirLLM 70B streaming on RTX 4060.
To associate your repository with the airllm topic, visit your repo's landing page and select "manage topics."