Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -12,9 +12,11 @@ Base change: PR #1596, `fix(codex): restore deferred tool discovery for non-Curs
> bundle was authored without a mounted checkout; every claim below was
> re-verified on 2026-08-13 against a real worktree, the upstream `codex-rs`
> source and live GitHub state. Eight corrections were recorded, the most
> consequential being that an **eligible** MCP tool stays callable in **both**
> discovery modes under code mode — so for those tools `direct` is a
> comprehension/compatibility lever with a payload cost, not a reachability fix.
> consequential being that a source reading indicates an **eligible** MCP tool stays
> callable in **both** discovery modes under code mode — so for those tools `direct`
> reads as a comprehension/compatibility lever with a payload cost rather than a
> reachability fix. That conclusion is not yet backed by an executed single-variable
> differential; `020` still owes it and `029` records it as open.
> "Eligible" excludes `direct_only_tool_namespaces`, `excluded_tool_namespaces`,
> and anything removed by MCP/App policy filtering; see `094` for the exclusion
> table and the differential test that must prove the claim.
Expand Down Expand Up @@ -48,7 +50,9 @@ At the verified `dev` head:

### PR A — profile resolver and explicit escape hatch

Default behavior remains byte-for-byte equivalent to #1596:
Default behavior is intended to be equivalent to #1596, and is asserted per-key on both
construction paths rather than by a normalized diff against a real prior build (that
comparison is still open — see `029`):

- non-Cursor: deferred
- Cursor: direct
Expand Down
Original file line number Diff line number Diff line change
@@ -1,5 +1,37 @@
# 029 - Phase 2 exit gate

> **Status 2026-08-13: Phase 2 is NOT closed.** The configuration and
> catalog-policy half shipped (see PR "feat(codex): add routed tool-discovery
> compatibility profiles"), and an earlier revision of this note claimed the
> `020` differential was the only thing left. An independent audit disproved
> that. The honest remaining list:
>
> - **`020` single-variable code-mode differential** — still owed. Until it
> runs, the claim in `004`/`094` that `direct` is a comprehension lever rather
> than a reachability fix rests on a source reading of a 2026-07-23 upstream
> clone. It needs a running Codex client, so it belongs to the live phase
> rather than the configuration PR. The docs now label that conclusion
> source-derived rather than proven.
> - **`023` zero-config comparison against a real prior build** — the suite pins
> the exact emitted key set on both construction paths, which is strong
> regression coverage, but it is not the documented normalized diff of a
> current-`dev` catalog against a patched one.
> - **`025` forcing-member diagnostic** — the combo explain surface does not yet
> name which member forced `direct`. Derivation itself is covered.
>
> Closed since that earlier revision, with ablation evidence recorded in
> `tests/codex-tool-discovery-mode.test.ts`: the `024` concurrency set (a
> latched provider `fetch` proves identical policies JOIN one flight and
> differing policies SPLIT, including model-map divergence and warm-cache
> re-resolution), the full `025` five-row matrix plus deferred-key omission and
> every alias shape (bare, slashed, native), each pinned to an exact emitted slug so
> deleting alias propagation turns them red, the `023` on-disk load/save
> round trip and downgrade preservation, and the `020` malformed-load warning, including the blank-model-key case where the
> schema drops the entire map.
>
> The retracted claim is left visible on purpose: this unit exists because a
> plan asserted more verification than it had.

Phase 2 is complete only when:

- [ ] pure resolver tests pass;
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -157,9 +157,10 @@ executed differential rather than on a source reading. The expected delta is
`exec.description` content/size **plus** `tool_search` construction and the
deferred-guidance text — not `exec.description` alone.

This reframes the override honestly: for an eligible tool under code mode it is a
*model-comprehension* lever (the model sees full schemas inline instead of having
to consult `ALL_TOOLS`), not a *reachability* lever. `010`'s non-goal list already
This reframes the override honestly: on the source reading above, for an eligible tool
under code mode it is a *model-comprehension* lever (the model sees full schemas inline
instead of having to consult `ALL_TOOLS`) rather than a *reachability* lever — still
pending the executed differential `020` owes. `010`'s non-goal list already
says "no claim that `direct` is cheaper or preferred"; it must also say direct is
not a reachability fix for eligible tools under code mode. `044`'s weak-model
fallback rationale is the honest use case.
Expand Down
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
4e5dd4520224a37fddfd8fb0dc532cd62c9daa9f703d2037c2875a44f8073b58 ./000_master_plan.md
d6538ee7c9a522a99746ee860e7f6d0e4a98e4a7baeffbba78d57f00fafe0295 ./000_master_plan.md
7e853a53e344412cc4aa2fcbe64b140195e9825a1988c5e245204c19a8f22449 ./001_verified_dev_baseline.md
32d7a7ffff41bd0c036cef153f53047ce416624032dae9dff43e88ff24d653fc ./002_incident_history_1522_1529_1596.md
343ee877dde67978b7ffd6a1e6f7825751c5fb113fd51dead6554036f92fa904 ./003_current_code_map.md
Expand All @@ -20,7 +20,7 @@ a8ea01e49ccb8c964c56581455e45288651d84a13f12a1ad9fcbdfad41b8641f ./021_catalog_
bd18baa3e16b54fe262435a6ee28c1ca5b389ec8b54433be1b05466559df58b1 ./023_backward_compatibility_tests.md
9542600beb52c958f6065a2c76a18c104c791c01b2035860b9730296ae45b4de ./024_catalog_cache_identity_tests.md
4fdb03fa022bf684257f95752e545bd4e9541b9258361bd9f5da5fc137b822cc ./025_combo_policy_tests.md
00c0913b15591630290c4a08429393570467cdc6799a1433b0eb16532b0630b0 ./029_phase2_exit_gate.md
ac0cd8816fcf2f032d02798f25383220803084d2c3642f16eafa0c1cc2d13e37 ./029_phase2_exit_gate.md
1f49005dbac81b2c3f3548d97f9aa32d6c8ada99f533d0fe197bcb047c553464 ./030_phase3_protocol_conformance.md
5e450a7557f7be2fad3f06c3c35288eb8654f156c402f826ef9c9b9703a52bf6 ./031_responses_lite_additional_tools.md
7bc16b20b8abe848726cb662865d542f2a8a19d6acb400a5473b0e9e3bc003e1 ./032_custom_namespace_roundtrip.md
Expand Down Expand Up @@ -63,7 +63,7 @@ e4d78f0db15cfe7819c98c5fde166c863a74d3a751ffe8196ddb6492691dc799 ./083_incident
28cb35f98ab0827a18f7ab52f28ad8cfd3f21ef475142c38ff8d49d41b43c106 ./091_pr_stack_and_commits.md
5bf3d16bea1a3ef12d68e3a467b271d3fd6631202fa43fb5f862f4ed505a43a3 ./092_definition_of_done.md
4cd63329dc69b6af90d2bdd31e64926ebc4acc916e1b34ccbef66e44f9a10937 ./093_execution_order.md
d68e52650b5b4dd7a6047b77f7942148f9ee36414a193828847a65b239411124 ./094_landing_verification_pass.md
6923be79563be75d1933f1dde0aadd3998ba8387d3ca2f42621fbabc1ef28ef9 ./094_landing_verification_pass.md
383ed01bca1fda60f5b57fa13e0f9cc8408d1ed02a8eab802600f6afe4097c73 ./KOREAN_SUMMARY.md
61be1a43ef84b272e6d04170dbd3e0a617aac36d9d3fb0fc2c4d96da2a62f775 ./README.md
82522bf74dd85254d7162ce103886b2a483d1c274a241afb80fa488a8a7d4356 ./patches/0001-add-tool-discovery-module.patch
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,8 @@ account を削除しても mapping は保持され、同じ id を再追加す
| `noPenaltyModels?` | `string[]` |存在/周波数ペナルティを拒否するモデル。 |
| `noStructuredOutputModels?` | `string[]` | `openai-chat` エンドポイントが `response_format` を拒否する正確なモデル ID。要求モデルが項目と完全一致する場合だけフィールドを省略し、その他の `openai-chat` モデルでは structured-output 変換を維持します。 |
| `parallelToolCalls?` | `boolean` |並列ツール呼び出しを切り替えます。 OpenAI Chat はデフォルトでオンになっています。非チャット アダプターは明示的な `true` でのみアドバタイズします。 |
| `routedToolDiscovery?` | `"auto" \| "deferred" \| "direct"` | ルーティング行の Codex ツール探索ポリシーです。`auto`(デフォルト)は出荷時の挙動を保ちます。非 Cursor のルーティング行は deferred、Cursor は direct で、Cursor はハードフェンスのため `deferred` を設定しても無視されます。deferred 探索と非互換だと実証された経路にのみ `direct` を使ってください。最初のリクエストに MCP 宣言がすべて載って turn-1 ペイロードが実測で 2.7 倍になり、コードモードでは Codex がどちらでも `tools`/`ALL_TOOLS` グローバルに入れ子ツールを設置するため、本来利用可能なツールの到達性が回復するわけでは**なさそうです**(上流 `codex-rs` のソースを読んだ判断であり、変数を 1 つだけ変えた実行比較ではまだ確認していません)。ホスト型ウェブ検索(`web_search_tool_type`)とは独立です。 |
| `modelRoutedToolDiscovery?` | `Record<string, "auto" \| "deferred" \| "direct">` | `routedToolDiscovery` のモデル単位の上書きです。混在ゲートウェイで非互換な 1 モデルのために兄弟モデルが不利にならないようにします。マッチングは通常のモデルキー規則(完全一致の id、`:` より前のファミリー、大文字小文字を無視)に従い、日付付きの `-YYYYMMDD` 変種はマッチしないため、失敗したモデル id を正確に指定してください。 |
| `responsesItemIdRepair?` | `{ message?: string[]; reasoning?: string[]; repairMissingTerminalIds?: boolean; repairInvalidIds?: boolean }` |正確なプレースホルダー ID、欠落している端末 ID、および(`repairInvalidIds` で)正規の `msg_`/`rs_` 接頭辞を欠く message/reasoning ID に対するダウンストリーム SSE 修復はデフォルトで無効になっています。関数呼び出し ID は決して書き換えられません。組み込み DeepSeek は最後の 2 つをデフォルトで有効にします。 |
| `responsesSnapshotRepair?` | `boolean` | デフォルトで無効のクライアント向け修復です。SSE と JSON の Responses ライフサイクルで欠落した status、output、ツールメタデータを補完し、raw 検査と永続化は変更しません。 |
| `retryOn429?` | `{ enabled?: boolean; attempts?: number; intervalMs?: number; maxIntervalMs?: number; respectRetryAfter?: boolean }` | API-key プロバイダーのみ(`authMode: "key"`)。オプトインの同一ターゲット 429 リトライ: `retryOn429` が無ければ無効で、オブジェクトがあれば `enabled: false` でない限り有効になります。429 時に待機(上流の `Retry-After` または固定間隔)してから、キー フェイルオーバーの前に同一キーで同一リクエストを再送します — メインのテキストターン回復ループ、Responses passthrough、画像/動画ブリッジ、web-search サイドカー、ターミナル継続要求をすべてカバーします。再送の対象はプリストリームの HTTP 429 応答のみで、カスタム `runTurn` トランスポートは HTTP リトライループの対象外です。`attempts` は最初の 429 以降の同一キー再送回数(合計送信数 = `attempts` + 1)で、メインの回復ループ・ターミナルガード継続・ブリッジ再試行で共有されるリクエスト単位の予算です。`attempts` を使い切っても同一キーでの再送が止まるだけで、通常のキー フェイルオーバーまたは最終エラー処理が利用可能なターゲットに応じて続きます — キー認証の passthrough ワイヤにはフェイルオーバーがないため、使い切った 429 はそのまま返ります。Codex 自体は 429 をリトライしないため、単一キーのプロバイダーでは唯一の防御です。デフォルト: `enabled: true`、`attempts: 3`、`intervalMs: 5000`、`maxIntervalMs: 60000`(1回の待機は `maxIntervalMs` で上限、その上限は 600000)、`respectRetryAfter: true`。 |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,8 @@ managed map을 활성화하면 privacy-safe selector를 만들고, 이후 계정
| `noPenaltyModels?` | `string[]` | presence/frequency penalty를 허용하지 않는 모델입니다. |
| `noStructuredOutputModels?` | `string[]` | `openai-chat` 엔드포인트가 `response_format`을 거부하는 정확한 모델 ID입니다. 요청 모델이 항목과 정확히 일치할 때만 필드를 생략하며, 그 외 `openai-chat` 모델에서는 structured-output 변환을 유지합니다. |
| `parallelToolCalls?` | `boolean` | 병렬 도구 호출을 켜거나 끕니다. OpenAI Chat은 기본으로 켜져 있고, 비-chat 어댑터는 명시적으로 `true`일 때만 이를 노출합니다. |
| `routedToolDiscovery?` | `"auto" \| "deferred" \| "direct"` | 라우팅된 행의 Codex 도구 탐색 정책입니다. `auto`(기본값)는 출하된 동작을 그대로 유지합니다. 비-Cursor 라우팅 행은 deferred, Cursor는 direct이며 Cursor는 하드 펜스라 `deferred`를 설정해도 무시합니다. deferred 탐색과 맞지 않는다고 입증된 경로에만 `direct`를 쓰십시오. 첫 요청에 모든 MCP 선언이 실려 turn-1 페이로드가 측정상 2.7배로 늘고, 코드 모드에서는 Codex가 어느 쪽이든 `tools`/`ALL_TOOLS` 전역에 중첩 도구를 심기 때문에 자격을 갖춘 도구의 도달성을 되살려 주지는 **않는 것으로 보입니다**(업스트림 `codex-rs` 소스를 읽어 판단했고, 변수를 하나만 바꾼 실행 대조 실험으로는 아직 확인하지 않았습니다). 호스티드 웹 검색(`web_search_tool_type`)과는 무관합니다. |
| `modelRoutedToolDiscovery?` | `Record<string, "auto" \| "deferred" \| "direct">` | `routedToolDiscovery`의 모델별 재정의입니다. 혼합 게이트웨이에서 호환되지 않는 한 모델 때문에 형제 모델까지 손해 보지 않게 합니다. 매칭은 기존 모델 키 규칙(정확한 id, `:` 앞 계열, 대소문자 무시)을 따르며 날짜가 붙은 `-YYYYMMDD` 변형은 매칭하지 않으므로 실패한 모델 id를 정확히 적으십시오. |
| `responsesItemIdRepair?` | `{ message?: string[]; reasoning?: string[]; repairMissingTerminalIds?: boolean; repairInvalidIds?: boolean }` | 기본값이 꺼진 downstream SSE 복구입니다. 정확한 자리표시자 id, 누락된 종료 id, 그리고(`repairInvalidIds`) 정규 `msg_`/`rs_` 접두사가 없는 message/reasoning id를 복구합니다. function-call id는 다시 쓰지 않습니다. 내장 DeepSeek은 마지막 두 가지를 기본으로 켭니다. |
| `responsesSnapshotRepair?` | `boolean` | 기본값이 꺼진 클라이언트용 복구입니다. SSE와 JSON의 Responses 수명 주기에서 누락된 status, output, 도구 메타데이터를 채우며 raw 검사와 영속화는 변경하지 않습니다. |
| `retryOn429?` | `{ enabled?: boolean; attempts?: number; intervalMs?: number; maxIntervalMs?: number; respectRetryAfter?: boolean }` | API-key 프로바이더 전용(`authMode: "key"`). 동일 대상 429 재시도: `retryOn429`가 없으면 기능이 꺼져 있고, 객체가 있으면 `enabled: false`가 아닌 한 활성화됩니다. 429 시 대기(업스트림 `Retry-After` 또는 고정 간격) 후 키 장애 조치 전에 동일 키로 동일 요청을 재전송합니다 — 일반 텍스트 턴 복구 루프, Responses passthrough, 이미지/비디오 브리지, web-search 사이드카, 터미널 연속 요청을 모두 포함합니다. 재전송 대상은 프리스트림 HTTP 429 응답뿐이며, 커스텀 `runTurn` 전송은 HTTP 재시도 루프에서 제외됩니다. `attempts`는 첫 429 이후의 동일 키 재전송 횟수(총 전송 = `attempts` + 1)이며, 메인 복구 루프·터미널 가드 연속 요청·브리지 재시도가 공유하는 요청 단위 예산입니다. `attempts`를 모두 소진해도 동일 키 재전송만 중단되며, 이후에는 일반 키 장애 조치 또는 최종 오류 처리가 사용 가능한 대상에 따라 진행됩니다 — 키 인증 passthrough 와이어에는 장애 조치가 없으므로 소진된 429가 그대로 반환됩니다. Codex 자체는 429를 재시도하지 않으므로 단일 키 프로바이더의 유일한 방어선입니다. 기본값: `enabled: true`, `attempts: 3`, `intervalMs: 5000`, `maxIntervalMs: 60000`(단일 대기는 `maxIntervalMs`로 상한, 그 자체는 600000으로 상한), `respectRetryAfter: true`. |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,8 @@ differing backup and rewrites known legacy namespaced selected ids to bare ids.
| `noPenaltyModels?` | `string[]` | Models that reject presence/frequency penalties. |
| `noStructuredOutputModels?` | `string[]` | Exact model IDs whose `openai-chat` endpoint rejects `response_format`. Only an exact requested-model match omits the field; structured-output translation stays enabled for every other `openai-chat` model. |
| `parallelToolCalls?` | `boolean` | Toggle parallel tool calls. OpenAI Chat defaults on; non-chat adapters advertise only on explicit `true`. |
| `routedToolDiscovery?` | `"auto" \| "deferred" \| "direct"` | Routed-row Codex tool-discovery policy. `auto` (default) keeps the shipped behavior: deferred for non-Cursor routed rows, direct for Cursor, which is hard-fenced and ignores a configured `deferred`. Set `direct` only for a route proven incompatible with deferred discovery — it embeds every MCP declaration in the first request (a measured 2.7x turn-1 payload cost) and does **not** appear to make an otherwise eligible tool reachable under code mode, where Codex installs nested tools on the `tools`/`ALL_TOOLS` globals either way (read from the upstream `codex-rs` source, not yet confirmed by an executed single-variable differential). Independent of hosted web search (`web_search_tool_type`). |
| `modelRoutedToolDiscovery?` | `Record<string, "auto" \| "deferred" \| "direct">` | Per-model override of `routedToolDiscovery`, so one incompatible model on a mixed gateway does not penalize its siblings. Matching follows the usual model-key rules (exact id, family before `:`, case-insensitive); dated `-YYYYMMDD` variants are not matched, so name the exact failing model id. |
| `responsesItemIdRepair?` | `{ message?: string[]; reasoning?: string[]; repairMissingTerminalIds?: boolean; repairInvalidIds?: boolean }` | Disabled-by-default downstream SSE repair for exact placeholder ids, missing terminal ids, and (with `repairInvalidIds`) message/reasoning ids missing the canonical `msg_`/`rs_` prefix. Function-call ids are never rewritten. Built-in DeepSeek enables the last two by default. |
| `responsesSnapshotRepair?` | `boolean` | Disabled-by-default client-facing repair for sparse Responses lifecycle snapshots in SSE and JSON. Fills missing canonical status, output, and tool metadata while raw inspection and persistence remain unchanged. |
| `retryOn429?` | `{ enabled?: boolean; attempts?: number; intervalMs?: number; maxIntervalMs?: number; respectRetryAfter?: boolean }` | API-key providers only (`authMode: "key"`). Opt-in same-target 429 retry: when `retryOn429` is absent the feature is off; object presence enables it unless `enabled: false`. On 429 the proxy waits (upstream `Retry-After` or the fixed interval) and replays the identical request on the same key before any key failover — across the main text-turn recovery loop, the Responses passthrough wire, the image/video bridge, the web-search sidecar, and terminal continuations. Only pre-stream HTTP 429 responses are eligible for replay; custom `runTurn` transports are outside the HTTP retry loop. `attempts` counts same-key replays after the first 429 (total sends = `attempts` + 1) and is one request-wide budget shared by the main recovery loop, the terminal-guard continuation, and bridge retries. Exhausting `attempts` only stops further same-key replays: normal key failover or final-error handling then applies per the available targets — on the key-auth passthrough wire there is no failover, so the exhausted 429 surfaces as-is. Codex itself never retries 429, so this is the only defense for single-key providers. Defaults: `enabled: true`, `attempts: 3`, `intervalMs: 5000`, `maxIntervalMs: 60000` (any single wait is capped at `maxIntervalMs`, itself capped at 600000), `respectRetryAfter: true`. |
Expand Down
Loading
Loading