Skip to content

[Bug]: ChatGPT codex backend 400s on gpt-5.6 models whenever prompt_cache_retention is forwarded (and intermittently even without it) #2092

Description

@pumpkinredbean

Client or integration

Codex App

Area

Provider adapter

Summary

Requests to gpt-5.6-luna / gpt-5.6-sol through the openai provider (ChatGPT codex backend) fail with 400 whenever the inbound request carries prompt_cache_retention. The remote Codex App started emitting this parameter on some turn requests (the string does not exist anywhere in codex-rs 0.146.0 — it is injected client-side), opencodex forwards request bodies opaquely, and one 400 aborts the whole agent turn mid-run.

Deterministic reproduction on stock 2.25.0 (proxy at localhost:10100, provider openaihttps://chatgpt.com/backend-api/codex, authMode forward, account pool):

# prompt_cache_retention: "24h"      → 400 {"detail":"Unsupported parameter: prompt_cache_retention"}  (3/3)
# prompt_cache_retention: "in_memory" → 400 {"detail":"Unsupported parameter: prompt_cache_retention"}  (3/3)
# parameter omitted                   → 200 OK  (3/3)
# client-shaped (service_tier=priority, prompt_cache_key, client_metadata, reasoning.effort=max,
#   parameter omitted)                → 200 OK  (3/3)

Request-history confirms the 400s are real upstream round trips (~300-600ms, sendCount: 1, account main), not proxy-side validation.

Additional upstream inconsistency observed the same day on 2.11.1: the same error class also fired intermittently on requests whose wire body provably did not contain the parameter (verified by logging the serialized outbound body at the failure point: prompt_cache_retention=undefined, 21 failures / ~9,100 requests over ~2h on one busy session, clustered around compaction-heavy periods). Earlier that morning one deployment accepted "24h" (response echoed prompt_cache_retention:"24h"), while another rejected the same value allowlist-style hours later. So the backend appears to apply account-level cache defaults with per-deployment disagreements; the one invariant is that the parameter must never be sent.

Reproduction

  1. ocx start --port 10100 (provider openai → ChatGPT codex backend, forward auth, pool mode)
  2. curl -N -X POST http://127.0.0.1:10100/v1/responses -H 'Content-Type: application/json' -d '{"model":"gpt-5.6-luna","input":[{"type":"message","role":"user","content":[{"type":"input_text","text":"say ok"}]}],"prompt_cache_retention":"24h","store":false,"stream":true}'
  3. Observe {"detail":"Unsupported parameter: prompt_cache_retention"} (400)

In production the parameter arrives from the Codex App itself: threads on gpt-5.6 models intermittently die with Provider error 400: {"error":{"message":"prompt_cache_retention is not supported on this model","code":"unsupported_api_for_model"-class errors}}.

Version

2.25.0 (reproduced); also observed on 2.11.1. No reference to prompt_cache_retention exists in the 2.25.0 source — bodies are passed through verbatim.

Operating system

macOS 15.5 (Apple Silicon)

Provider and model

openai (ChatGPT codex backend) / gpt-5.6-luna, gpt-5.6-sol

Logs or error output

# Upstream response (stock 2.25.0, request carrying the parameter):
{"detail":"Unsupported parameter: prompt_cache_retention"}

# Client-facing failure observed on 2.11.1 (turn aborts mid-run):
Provider error 400: {"error":{"message":"prompt_cache_retention is not supported on this model","type":"invalid_request_error","param":"prompt_cache_retention","code":"invalid_parameter"}}

# Outbound-body log at the moment of one such 400 (note: parameter absent from the wire):
wire prompt_cache_retention=undefined model=gpt-5.6-luna store=false
keys=model|input|tool_choice|parallel_tool_calls|reasoning|store|stream|stream_options|include|service_tier|prompt_cache_key|text|client_metadata

Redacted configuration

{
  "providers": {
    "openai": {
      "adapter": "openai-responses",
      "baseUrl": "https://chatgpt.com/backend-api/codex",
      "codexAccountMode": "pool",
      "authMode": "forward"
    }
  }
}

Suggested handling / local workaround

Strip the parameter in the openai-responses passthrough adapter so a client-emitted value can never reach the wire (the backend still applies its own account-level cache default — responses keep echoing prompt_cache_retention:"24h"). This mirrors the existing stripUnsupportedReasoningParams / issue-#323 precedents. I have this running locally with the reproduction matrix going 400 → 200 across the board; happy to open a PR against dev with a focused regression test.

+function stripUnsupportedPromptCacheRetention(body: unknown): unknown {
+  if (!isPlainObject(body)) return body;
+  if (!Object.hasOwn(body, "prompt_cache_retention")) return body;
+  const { prompt_cache_retention: _retention, ...rest } = body;
+  return rest;
+}

(wired into the outBody pipeline in createResponsesPassthroughAdapter.buildRequest)

Optionally also consider treating this specific 400 as one-time retryable: during the intermittent window, identical param-less requests alternately failed and succeeded within minutes.

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Metadata

Metadata

Assignees

No one assigned

    Labels

    account-poolOAuth, credentials, Codex pool, quota, failover, plansbugSomething isn't workingproviderProvider adapters, OpenAI-compat presets, upstream API quirks

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions