Allow Realtime voice conversation to call preview tool - #219
Merged
MikeAlhayek merged 1 commit intoSep 25, 2026
Merged
MikeAlhayek merged 1 commit into
MikeAlhayek merged 1 commit into
Conversation
A voice reply's text is a transcript of what the model said aloud, so it never contains the [fig:N] marker the client swaps for the preview picture. The preview was built and sent with the reply, but never drawn. The preview tool now asks the host to show each picture it registers, and the realtime runner appends the marker to the next spoken turn. The document prompt also sends preview requests to the tabular agent and tells a voice model not to read markers aloud.
Contributor
Author
|
fixes #220 |
Contributor
Author
|
@MikeAlhayek This is ready for review Tested and working on my end. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Asking for a preview of an uploaded Excel file in a realtime voice conversation (
gpt-realtime) never showed the picture. The same request typed into the same chat did. This makes the voice conversation show it too.Why it did not work on
mainThe voice model was calling the tool. In the failing session the log shows
tabular-data-agentrunningpreview_tabular_dataand registering[fig:1]for the generated SVG. The saved voice turns carried that image reference. But the browser never requested the picture: there is noDownloadAIDocumentline for it, where the typed turn earlier in the same chat has one. The picture was built and sent, and never drawn.[fig:N]and tells the model to write that marker in its answer (PreviewTabularDataTool.csL504-L516).chat-markers.jsL58-L96, called fromchat-interaction.jsL351).RealtimeChatSessionRunner.csL516-L530), and the live bubble is built from those tokens (chat-interaction.jsL880-L928). Nobody says "[fig:1]" aloud, so there was never a marker to replace. The saved reply read "Here's the image preview of the first 5 rows of the spreadsheet:" and ended there.ChatInteractionHubBase.csL1729-L1741,BuildUncitedImageMarkers). The realtime runner saves the transcript as it is (FlushAssistantTurnAsyncL939-L975).Two prompt gaps made it worse:
document-availability.mdL37). In a second voice session the model never delegated, and described the sheets fromget_document_metadatainstead.agent-availability.mdL23). A voice model can only do that by speaking it. It also calledview_document_figuretrying to put the picture on screen itself.What changed, and why it works now
AIInvocationContext.RequestFigureDisplay/TakeFigureDisplayRequestshold a thread-safe queue of markers. Tools run under the function-invocation middleware while the runner flushes turns from its own event pump, so the queue is concurrent.PreviewTabularDataTool.csL513-L519). It asks again when a repeat request is answered from its cache (L131-L138). A voice session is one invocation, so a second "show me the preview" minutes later hits that cache.RealtimeChatSessionRunner.csL959-L996).SnapshotReferencesonmain). Appending every uncited image would repeat the preview under every later reply, and would add retrieved figures the model never chose to show.document-availability.mdL40).view_document_figurefor one (L47-L50).isRealtimeargument (DocumentOrchestrationHandler.csL211), so typed chat is unaffected.Verification
RealtimeChatSessionRunnerTests.csL1324-L1524) cover:gpt-realtime:preview_tabular_dataregistered[fig:1], and the browser downloaded the SVG six seconds later. The saved reply reads "You should see the preview on your screen now. It shows a table with columns like client names, revenue projections… [fig:1]". The picture appears once, on that reply, and not under later turns.Not covered here
[doc:N]exports) probably build up across a voice session for the same reason as point 4, and may be listed under every later reply. Not changed here, and not yet confirmed.🤖 Generated with Claude Code