Skip to content

Finish the chat response before the session postcommit (opt-in) - #557

Open
jvmenen wants to merge 1 commit into
youssofal:mainfrom
jvmenen:pr/postcommit-after-response
Open

jvmenen wants to merge 1 commit into
youssofal:mainfrom
jvmenen:pr/postcommit-after-response

Conversation

@jvmenen

@jvmenen jvmenen commented Sep 28, 2026

Copy link
Copy Markdown

Summary

Opt-in MTPLX_POSTCOMMIT_AFTER_RESPONSE=1: a chat turn's terminal SSE frame (or the non-stream JSON body) is released before the generation-final session postcommit instead of after it. The commit runs as a post-response tail, and the next chat or completions request waits for every tail before it reads session or bank state, so no request ever sees a half-committed session. Response text is unchanged. Default off.

Motivation

At the end of a chat turn with a session, the server stores the generation-final history snapshot (re-render and encode of the history, then the bank write) before it sends the terminal frame or returns the body. The client sees the response end only after that work, although the tokens are all generated and nothing in the response depends on the commit's outcome. For a client that waits for the end of the response before it acts (an agent harness that runs the tool call, for example), that time adds to every turn.

The commit only has to land before the next request of that session looks at the bank. Ordering it after the response and gating admission on it keeps that guarantee.

Change

All in mtplx/server/openai.py:

  • _postcommit_after_response_enabled() (MTPLX_POSTCOMMIT_AFTER_RESPONSE, default off) and _post_response_tail_wait_s() (MTPLX_POSTCOMMIT_AFTER_RESPONSE_WAIT_S, default max(60, STREAM_COMMIT_WAIT_MAX_S + 30)): an upper bound for one wait, so a wedged tail cannot block admission for good.
  • _PostResponseTails: counts and tracks the running tails on the server state; wait_idle(timeout) blocks until none are running.
  • Streaming: with the switch on, the stream worker releases the terminal frame and hands the snapshot to a tail task, which waits for the worker to leave the generation slot and then runs the commit as serial model work. Non-stream: the JSON body is returned first and the same commit runs as a tail.
  • Admission barrier: /v1/chat/completions (and through it /v1/messages) and /v1/completions await _await_post_response_tails before resolving the session or looking up the bank; the wait is recorded in the request observability (post_response_tail_wait).
  • The response's session_postcommit_snapshot reads after_response; the real outcome is merged into that request's metrics row when the tail finishes (_merge_post_response_stats touches only its own row). /health reports post_response_tails.
  • With the switch off, the code paths are the existing ones. Most of the diff is re-indentation of the existing commit code into the tail functions; git diff -w shows the substance.

How to enable: MTPLX_POSTCOMMIT_AFTER_RESPONSE=1.

Tests

  • New tests/test_postcommit_after_response.py (10 tests, no model, FastAPI test client with the fake streaming session from test_server_openai.py): flag and wait-bound parsing; wait_idle blocks until the tail leaves; the admission barrier's receipt; the metrics merge only touches its own row; streaming and non-stream: with the switch off the postcommit completes before the terminal frame or body, with it on the response ends first and the next turn waits for the commit; the text is identical with and without the switch.
  • tests/test_server_openai.py and tests/test_postcommit_prefix_reuse.py pass.
  • Full suite with MTPLX_CONFIG=/nonexistent: 9,391 passed, 67 skipped, 1 failed. The failure is test_laguna_model.py::test_laguna_s_2_1_ar_route_skips_qwen_performance_hooks, which fails the same way on main on a Mac with less than 85.3 GiB (addressed in Tests no longer read the user's ~/.mtplx/config.toml or depend on the Mac's memory size #535). Ruff: no new findings.

🤖 Generated with Claude Code

…mit (opt-in)

MTPLX_POSTCOMMIT_AFTER_RESPONSE=1 (default off) releases the terminal SSE
frame or the JSON body before the generation-final history snapshot, whose
history re-encode otherwise delays the end of the response. The commit runs
as a post-response tail; chat and completions admission waits for every
tail (and for the stream worker to leave the generation slot) before it
reads session or bank state, so a next turn never sees a half-committed
session. Text output is unchanged; the response's
session_postcommit_snapshot reads after_response and the real outcome
lands in the request metrics. /health reports post_response_tails.
MTPLX_POSTCOMMIT_AFTER_RESPONSE_WAIT_S bounds one wait on the tails.

Tests: tests/test_postcommit_after_response.py: flag and wait-bound
parsing, the tail tracker and admission barrier, metrics merge; stream and
non-stream: off waits for the postcommit before the terminal frame, on
ends the response first and the next turn waits for the commit, and the
text is identical with and without the flag.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jvmenen
jvmenen requested a review from youssofal as a code owner September 28, 2026 09:50

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant