Skip to content

Allow Realtime voice conversation to call preview tool - #219

Merged
MikeAlhayek merged 1 commit into
CrestApps:mainfrom
JackTelford:jt-allow-real-time-build-preview-excel-file
Sep 25, 2026
Merged

MikeAlhayek merged 1 commit into
CrestApps:mainfrom
JackTelford:jt-allow-real-time-build-preview-excel-file

Conversation

@JackTelford

Copy link
Copy Markdown
Contributor

Asking for a preview of an uploaded Excel file in a realtime voice conversation (gpt-realtime) never showed the picture. The same request typed into the same chat did. This makes the voice conversation show it too.

Why it did not work on main

The voice model was calling the tool. In the failing session the log shows tabular-data-agent running preview_tabular_data and registering [fig:1] for the generated SVG. The saved voice turns carried that image reference. But the browser never requested the picture: there is no DownloadAIDocument line for it, where the typed turn earlier in the same chat has one. The picture was built and sent, and never drawn.

  1. The preview is a marker the model must write. The tool registers the picture under [fig:N] and tells the model to write that marker in its answer (PreviewTabularDataTool.cs L504-L516).
  2. The client only draws a picture where its marker appears in the reply text (chat-markers.js L58-L96, called from chat-interaction.js L351).
  3. A voice reply's text is only what the model said aloud. The runner builds the reply from the spoken transcript (RealtimeChatSessionRunner.cs L516-L530), and the live bubble is built from those tokens (chat-interaction.js L880-L928). Nobody says "[fig:1]" aloud, so there was never a marker to replace. The saved reply read "Here's the image preview of the first 5 rows of the spreadsheet:" and ended there.
  4. Typed chat has a safety net that voice did not. When a typed answer leaves out an image marker, the hub appends it (ChatInteractionHubBase.cs L1729-L1741, BuildUncitedImageMarkers). The realtime runner saves the transcript as it is (FlushAssistantTurnAsync L939-L975).

Two prompt gaps made it worse:

  • Previews were not listed as a tabular-agent task. The delegation instruction lists summaries, filtering, calculations and so on, but never showing or previewing the file (document-availability.md L37). In a second voice session the model never delegated, and described the sheets from get_document_metadata instead.
  • The parent model is told to copy every marker into its answer (agent-availability.md L23). A voice model can only do that by speaking it. It also called view_document_figure trying to put the picture on screen itself.

What changed, and why it works now

  1. A tool can ask the host to show a picture. AIInvocationContext.RequestFigureDisplay / TakeFigureDisplayRequests hold a thread-safe queue of markers. Tools run under the function-invocation middleware while the runner flushes turns from its own event pump, so the queue is concurrent.
  2. The preview tool asks for each picture it registers (PreviewTabularDataTool.cs L513-L519). It asks again when a repeat request is answered from its cache (L131-L138). A voice session is one invocation, so a second "show me the preview" minutes later hits that cache.
  3. The realtime runner appends requested markers to the next spoken turn (RealtimeChatSessionRunner.cs L959-L996).
    • The marker goes out as a transcript delta, so the live bubble gets it, and is saved with the turn, so it survives a reload. The client then draws the picture exactly as it does for a typed answer, so no client changes are needed.
    • It is never placed on a tool-only response. A marker the reply already contains, or one without a servable image, is skipped.
    • If the session ends before the model speaks again, the picture is saved as its own turn (L702).
  4. Why an explicit request instead of the typed path's "append every uncited image". A realtime invocation scope lasts the whole conversation, so references pile up across turns (SnapshotReferences on main). Appending every uncited image would repeat the preview under every later reply, and would add retrieved figures the model never chose to show.
  5. Prompt:
    • Preview requests now go to the tabular agent (document-availability.md L40).
    • A voice session is told the picture appears on screen automatically, never to read a marker aloud, and not to call view_document_figure for one (L47-L50).
    • The voice paragraph is gated on the new isRealtime argument (DocumentOrchestrationHandler.cs L211), so typed chat is unaffected.

Verification

  • 5 new runner tests (RealtimeChatSessionRunnerTests.cs L1324-L1524) cover:
    • placing a picture on the next spoken reply, not on the tool-only response
    • not duplicating a marker the reply already contains
    • showing it again on a repeat request
    • saving it as its own turn when the session ends first
    • ignoring unrequested, unservable and unregistered pictures
  • Full suite: 4016 passed, 0 failed.
  • Typed regression check: "Show me the preview of this file." still delegates and draws the picture.
  • Live voice test with gpt-realtime: preview_tabular_data registered [fig:1], and the browser downloaded the SVG six seconds later. The saved reply reads "You should see the preview on your screen now. It shows a table with columns like client names, revenue projections… [fig:1]". The picture appears once, on that reply, and not under later turns.

Not covered here

  • A repeat preview request within the same voice session is covered by a unit test but has not been exercised live yet.
  • Generated downloads ([doc:N] exports) probably build up across a voice session for the same reason as point 4, and may be listed under every later reply. Not changed here, and not yet confirmed.

🤖 Generated with Claude Code

A voice reply's text is a transcript of what the model said aloud, so it never
contains the [fig:N] marker the client swaps for the preview picture. The
preview was built and sent with the reply, but never drawn.

The preview tool now asks the host to show each picture it registers, and the
realtime runner appends the marker to the next spoken turn. The document prompt
also sends preview requests to the tabular agent and tells a voice model not to
read markers aloud.
@JackTelford

Copy link
Copy Markdown
Contributor Author

fixes #220

@JackTelford JackTelford changed the title Show the spreadsheet preview in a realtime voice conversation Allow Realtime voice conversation to call excel preview tool Sep 24, 2026
@JackTelford JackTelford changed the title Allow Realtime voice conversation to call excel preview tool Allow Realtime voice conversation to call preview tool Sep 24, 2026
@JackTelford

Copy link
Copy Markdown
Contributor Author

@MikeAlhayek This is ready for review Tested and working on my end.

@MikeAlhayek
MikeAlhayek merged commit e8c3684 into CrestApps:main Sep 25, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants