A compact local AI API proxy with protocol passthrough, retries, model routing, request adaptation, and single-binary deployment.
Many AI coding clients speak different wire protocols even when they ultimately target similar model providers.
mini-proxy gives them a single local base URL and routes requests by path:
OpenAI-style/chat/completions |
Anthropic-style/v1/messages |
Responses-style/responses |
It is designed for local model gateways, coding-plan APIs, compatibility layers, and setups where you want one stable endpoint in front of several upstream services.
| Capability | What it does |
|---|---|
| 🔀 Three protocol paths | OpenAI Chat Completions, Anthropic Messages, and Responses API style traffic |
| ♻️ Automatic retry | Retries by HTTP status range and provider business error code |
| 🌊 SSE streaming | Preserves streaming responses and can detect retryable errors during early stream events |
| 🔑 Two API-key modes | Client-key passthrough or configuration override |
| 🗺️ Model mapping | Maps client-facing model IDs to upstream model IDs |
| 🧹 Request cleanup | Optional semantic-aware cleanup; preserves tool/function messages even when text content is blank |
| 🧠 Reasoning-effort injection | Can force protocol-specific reasoning effort or leave client values untouched |
| 🧩 Provider inheritance | Shared provider settings with per-endpoint overrides |
| 🪪 Header injection | Adds configured headers only when the client did not already provide them |
| 🧵 OpenCode session support | Optional automatic x-opencode-session generation / propagation |
| 📦 Single executable | Build once, run directly; first launch can generate configuration |
| 🪵 Structured logging | Console + rolling file logs through tracing |
Download or build mini-proxy.exe, then run:
mini-proxy.exeOn first launch, mini-proxy can create config.toml and start with the configured defaults.
cargo build --releaseBinary:
target/release/mini-proxy
On Windows:
target/release/mini-proxy.exe
mini-proxy --helpCursor / Claude Code / Codex / other clients
│
▼
http://127.0.0.1:7946
│
┌───────────┼───────────┐
▼ ▼ ▼
/chat/completions /v1/messages /responses
OpenAI Anthropic Responses
│ │ │
└──── provider routing ──┘
│
model map / headers
retry / key handling
reasoning adaptation
│
▼
upstream APIs
| Protocol | Request path | SDK base URL |
|---|---|---|
| OpenAI-style | POST http://<listen>/chat/completions |
http://<listen> |
| Anthropic-style | POST http://<listen>/v1/messages |
http://<listen> |
| Responses-style | POST http://<listen>/responses |
http://<listen> |
Default listen address:
127.0.0.1:7946
The three protocols share the same base URL. mini-proxy distinguishes them by the incoming request path.
A minimal example:
[server]
listen = "127.0.0.1:7946"
clean_empty_content = false
[log]
level = "info"
format = "pretty"
to_stdout = true
to_file = "logs/proxy.log"
rotate_size_mb = 50
rotate_keep = 7
[[provider]]
name = "ExampleProvider"
api_key = ""
models = ["model-a", "model-b"]
max_retries = 10000
key_mode = "passthrough"
thinking_effort = "passthrough"
[provider.openai]
base_url = "https://example.com/v1"
[provider.anthropic]
base_url = "https://example.com"
[provider.responses]
base_url = "https://example.com/v1"The repository's config.toml contains a much more complete field reference and working provider examples.
| Field | Meaning | Typical / default behavior |
|---|---|---|
name |
Provider label used in logs | required |
api_key |
Key used in override mode | empty |
models |
Client-visible model IDs handled by this provider | provider-specific |
model_map |
Client model ID → upstream model ID | same-name passthrough when absent |
max_retries |
Maximum retries on the same provider/model | 10000 |
retry_on_status |
Retryable HTTP codes / ranges | built-in defaults when omitted |
retry_on_code |
Retryable provider business error codes | built-in defaults when omitted |
key_mode |
passthrough or override |
passthrough |
path_mode |
append or full |
append |
thinking_effort |
Force a reasoning level or pass the client value through | protocol-dependent default when omitted |
is_opencode |
Enable automatic OpenCode session header handling | false |
headers |
Static headers inserted only when absent from the request | empty |
Endpoint sections such as [provider.openai], [provider.anthropic], and [provider.responses] inherit provider-level values and may override them.
When retry_on_status is omitted, mini-proxy uses its built-in retry ranges. The current configuration documentation lists ranges covering most transient / provider-side failures while explicitly excluding 504 and 524 from retry.
mini-proxy can inspect error.code in the response body and retry selected provider-specific business errors.
For SSE responses, mini-proxy pre-reads the early stream events:
- a retryable
event: errorcan trigger a fresh upstream attempt; - once valid content begins, buffered events are forwarded together with the remainder of the stream.
This avoids committing a broken early stream to the client when the upstream failure is still recoverable.
| Mode | Behavior |
|---|---|
passthrough |
Preserve the client's key; mini-proxy does not need to own it |
override |
Replace the client key with api_key from configuration; an empty configured key falls back to passthrough |
| Mode | Behavior |
|---|---|
append |
base_url + protocol suffix, such as /chat/completions |
full |
Use base_url exactly as configured |
This is useful when different providers expose either a common API root or a fully specified endpoint.
Override OpenAI Base URL: http://127.0.0.1:7946
OpenAI API Key: <your key>
Model: <configured model id>
{
"env": {
"ANTHROPIC_AUTH_TOKEN": "<your key>",
"ANTHROPIC_BASE_URL": "http://127.0.0.1:7946",
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1",
"API_TIMEOUT_MS": "600000",
"ANTHROPIC_MODEL": "<configured model id>"
}
}auth.json:
{
"OPENAI_API_KEY": "<your key>"
}Example client configuration:
model_provider = "mini-proxy"
model = "<configured model id>"
disable_response_storage = true
preferred_auth_method = "apikey"
[model_providers.mini-proxy]
name = "mini-proxy"
base_url = "http://127.0.0.1:7946"
wire_api = "responses"Example:
[log]
level = "info"
format = "pretty"
to_stdout = true
to_file = "logs/proxy.log"
rotate_size_mb = 50
rotate_keep = 7The project uses tracing / tracing-subscriber and supports console output plus rolling file logs.
mini-proxy is currently a personal-use project. Configuration defaults and provider examples may evolve with the upstream services it is used against.
For Chinese documentation, see READMEch.md.