Conversation
…mit (opt-in) MTPLX_POSTCOMMIT_AFTER_RESPONSE=1 (default off) releases the terminal SSE frame or the JSON body before the generation-final history snapshot, whose history re-encode otherwise delays the end of the response. The commit runs as a post-response tail; chat and completions admission waits for every tail (and for the stream worker to leave the generation slot) before it reads session or bank state, so a next turn never sees a half-committed session. Text output is unchanged; the response's session_postcommit_snapshot reads after_response and the real outcome lands in the request metrics. /health reports post_response_tails. MTPLX_POSTCOMMIT_AFTER_RESPONSE_WAIT_S bounds one wait on the tails. Tests: tests/test_postcommit_after_response.py: flag and wait-bound parsing, the tail tracker and admission barrier, metrics merge; stream and non-stream: off waits for the postcommit before the terminal frame, on ends the response first and the next turn waits for the commit, and the text is identical with and without the flag. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Opt-in
MTPLX_POSTCOMMIT_AFTER_RESPONSE=1: a chat turn's terminal SSE frame (or the non-stream JSON body) is released before the generation-final session postcommit instead of after it. The commit runs as a post-response tail, and the next chat or completions request waits for every tail before it reads session or bank state, so no request ever sees a half-committed session. Response text is unchanged. Default off.Motivation
At the end of a chat turn with a session, the server stores the generation-final history snapshot (re-render and encode of the history, then the bank write) before it sends the terminal frame or returns the body. The client sees the response end only after that work, although the tokens are all generated and nothing in the response depends on the commit's outcome. For a client that waits for the end of the response before it acts (an agent harness that runs the tool call, for example), that time adds to every turn.
The commit only has to land before the next request of that session looks at the bank. Ordering it after the response and gating admission on it keeps that guarantee.
Change
All in
mtplx/server/openai.py:_postcommit_after_response_enabled()(MTPLX_POSTCOMMIT_AFTER_RESPONSE, default off) and_post_response_tail_wait_s()(MTPLX_POSTCOMMIT_AFTER_RESPONSE_WAIT_S, defaultmax(60, STREAM_COMMIT_WAIT_MAX_S + 30)): an upper bound for one wait, so a wedged tail cannot block admission for good._PostResponseTails: counts and tracks the running tails on the server state;wait_idle(timeout)blocks until none are running./v1/chat/completions(and through it/v1/messages) and/v1/completionsawait_await_post_response_tailsbefore resolving the session or looking up the bank; the wait is recorded in the request observability (post_response_tail_wait).session_postcommit_snapshotreadsafter_response; the real outcome is merged into that request's metrics row when the tail finishes (_merge_post_response_statstouches only its own row)./healthreportspost_response_tails.git diff -wshows the substance.How to enable:
MTPLX_POSTCOMMIT_AFTER_RESPONSE=1.Tests
tests/test_postcommit_after_response.py(10 tests, no model, FastAPI test client with the fake streaming session fromtest_server_openai.py): flag and wait-bound parsing;wait_idleblocks until the tail leaves; the admission barrier's receipt; the metrics merge only touches its own row; streaming and non-stream: with the switch off the postcommit completes before the terminal frame or body, with it on the response ends first and the next turn waits for the commit; the text is identical with and without the switch.tests/test_server_openai.pyandtests/test_postcommit_prefix_reuse.pypass.MTPLX_CONFIG=/nonexistent: 9,391 passed, 67 skipped, 1 failed. The failure istest_laguna_model.py::test_laguna_s_2_1_ar_route_skips_qwen_performance_hooks, which fails the same way onmainon a Mac with less than 85.3 GiB (addressed in Tests no longer read the user's ~/.mtplx/config.toml or depend on the Mac's memory size #535). Ruff: no new findings.🤖 Generated with Claude Code