Skip to content

About

Lightweight AI API proxy — OpenAI · Anthropic · Responses · SSE streaming · retries · model mapping / 轻量级 AI API 代理,支持流式转发、自动重试、模型映射与请求适配。

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

33 Commits

Folders and files

Repository files navigation

mini-proxy — one local endpoint for multiple AI API protocols

Chinese README Rust 2021 version 0.1.0 SSE streaming



A compact local AI API proxy with protocol passthrough, retries, model routing, request adaptation, and single-binary deployment.


Why mini-proxy?

Many AI coding clients speak different wire protocols even when they ultimately target similar model providers.

mini-proxy gives them a single local base URL and routes requests by path:

OpenAI-style
/chat/completions
Anthropic-style
/v1/messages
Responses-style
/responses

It is designed for local model gateways, coding-plan APIs, compatibility layers, and setups where you want one stable endpoint in front of several upstream services.


Highlights

Capability What it does
🔀 Three protocol paths OpenAI Chat Completions, Anthropic Messages, and Responses API style traffic
♻️ Automatic retry Retries by HTTP status range and provider business error code
🌊 SSE streaming Preserves streaming responses and can detect retryable errors during early stream events
🔑 Two API-key modes Client-key passthrough or configuration override
🗺️ Model mapping Maps client-facing model IDs to upstream model IDs
🧹 Request cleanup Optional semantic-aware cleanup; preserves tool/function messages even when text content is blank
🧠 Reasoning-effort injection Can force protocol-specific reasoning effort or leave client values untouched
🧩 Provider inheritance Shared provider settings with per-endpoint overrides
🪪 Header injection Adds configured headers only when the client did not already provide them
🧵 OpenCode session support Optional automatic x-opencode-session generation / propagation
📦 Single executable Build once, run directly; first launch can generate configuration
🪵 Structured logging Console + rolling file logs through tracing

Quick start

Windows

Download or build mini-proxy.exe, then run:

mini-proxy.exe

On first launch, mini-proxy can create config.toml and start with the configured defaults.

Build from source

cargo build --release

Binary:

target/release/mini-proxy

On Windows:

target/release/mini-proxy.exe

Inspect configuration help

mini-proxy --help

Request flow

Cursor / Claude Code / Codex / other clients
                    │
                    ▼
          http://127.0.0.1:7946
                    │
        ┌───────────┼───────────┐
        ▼           ▼           ▼
 /chat/completions /v1/messages /responses
      OpenAI        Anthropic    Responses
        │           │           │
        └──── provider routing ──┘
                    │
          model map / headers
        retry / key handling
        reasoning adaptation
                    │
                    ▼
             upstream APIs

Public endpoints

Protocol Request path SDK base URL
OpenAI-style POST http://<listen>/chat/completions http://<listen>
Anthropic-style POST http://<listen>/v1/messages http://<listen>
Responses-style POST http://<listen>/responses http://<listen>

Default listen address:

127.0.0.1:7946

The three protocols share the same base URL. mini-proxy distinguishes them by the incoming request path.


Configuration

A minimal example:

[server]
listen = "127.0.0.1:7946"
clean_empty_content = false

[log]
level = "info"
format = "pretty"
to_stdout = true
to_file = "logs/proxy.log"
rotate_size_mb = 50
rotate_keep = 7

[[provider]]
name = "ExampleProvider"
api_key = ""
models = ["model-a", "model-b"]
max_retries = 10000
key_mode = "passthrough"
thinking_effort = "passthrough"

[provider.openai]
base_url = "https://example.com/v1"

[provider.anthropic]
base_url = "https://example.com"

[provider.responses]
base_url = "https://example.com/v1"

The repository's config.toml contains a much more complete field reference and working provider examples.

Core provider fields

Field Meaning Typical / default behavior
name Provider label used in logs required
api_key Key used in override mode empty
models Client-visible model IDs handled by this provider provider-specific
model_map Client model ID → upstream model ID same-name passthrough when absent
max_retries Maximum retries on the same provider/model 10000
retry_on_status Retryable HTTP codes / ranges built-in defaults when omitted
retry_on_code Retryable provider business error codes built-in defaults when omitted
key_mode passthrough or override passthrough
path_mode append or full append
thinking_effort Force a reasoning level or pass the client value through protocol-dependent default when omitted
is_opencode Enable automatic OpenCode session header handling false
headers Static headers inserted only when absent from the request empty

Endpoint sections such as [provider.openai], [provider.anthropic], and [provider.responses] inherit provider-level values and may override them.


Retry behavior

HTTP status retry

When retry_on_status is omitted, mini-proxy uses its built-in retry ranges. The current configuration documentation lists ranges covering most transient / provider-side failures while explicitly excluding 504 and 524 from retry.

Provider error-code retry

mini-proxy can inspect error.code in the response body and retry selected provider-specific business errors.

Streaming retry

For SSE responses, mini-proxy pre-reads the early stream events:

  1. a retryable event: error can trigger a fresh upstream attempt;
  2. once valid content begins, buffered events are forwarded together with the remainder of the stream.

This avoids committing a broken early stream to the client when the upstream failure is still recoverable.


API key modes

Mode Behavior
passthrough Preserve the client's key; mini-proxy does not need to own it
override Replace the client key with api_key from configuration; an empty configured key falls back to passthrough

Upstream URL modes

Mode Behavior
append base_url + protocol suffix, such as /chat/completions
full Use base_url exactly as configured

This is useful when different providers expose either a common API root or a fully specified endpoint.


Client examples

Cursor — OpenAI-style endpoint

Override OpenAI Base URL: http://127.0.0.1:7946
OpenAI API Key: <your key>
Model: <configured model id>

Claude Code — Anthropic-style endpoint

{
  "env": {
    "ANTHROPIC_AUTH_TOKEN": "<your key>",
    "ANTHROPIC_BASE_URL": "http://127.0.0.1:7946",
    "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1",
    "API_TIMEOUT_MS": "600000",
    "ANTHROPIC_MODEL": "<configured model id>"
  }
}

Codex — Responses-style endpoint

auth.json:

{
  "OPENAI_API_KEY": "<your key>"
}

Example client configuration:

model_provider = "mini-proxy"
model = "<configured model id>"
disable_response_storage = true
preferred_auth_method = "apikey"

[model_providers.mini-proxy]
name = "mini-proxy"
base_url = "http://127.0.0.1:7946"
wire_api = "responses"

Logging

Example:

[log]
level = "info"
format = "pretty"
to_stdout = true
to_file = "logs/proxy.log"
rotate_size_mb = 50
rotate_keep = 7

The project uses tracing / tracing-subscriber and supports console output plus rolling file logs.


Tech stack


Project status

mini-proxy is currently a personal-use project. Configuration defaults and provider examples may evolve with the upstream services it is used against.

For Chinese documentation, see READMEch.md.

About

Lightweight AI API proxy — OpenAI · Anthropic · Responses · SSE streaming · retries · model mapping / 轻量级 AI API 代理,支持流式转发、自动重试、模型映射与请求适配。

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages