Design: typed text and realtime voice on the same chat UI - #214
Merged
MikeAlhayek merged 5 commits intoSep 21, 2026
Merged
Conversation
Today a realtime-capable chat deployment turns the whole chat UI voice-only, because one field answers two questions: AIProfile.ChatDeploymentName is both the text model and the realtime model. Users want to type and speak in the same thread, toggled on demand. This is design only. It works through what the coupling actually is (three places: the views' d-none, the ChatMode squash, and the hubs' text rejection), notes that the orchestration layer is already decoupled via RealtimeOrchestrationRequest.RealtimeDeploymentName and the realtime slot chain, and recommends a single new optional field -- the conversation deployment -- rather than storing a realtime-vs-STT/TTS transport choice that CascadedRealtimeMetadata already answers at the deployment level. Also records the one piece of real engineering: DefaultRealtimeOrchestrator starts realtime sessions with an empty ConversationHistory, which is invisible today but becomes the feature's most obvious bug once both modes share a thread.
A realtime-capable chat deployment turned the whole chat UI voice-only, because AIProfile.ChatDeploymentName was simultaneously the text model and the realtime model. Conversation mode now decides how the conversation is carried, and the chat deployment goes back to meaning only "the text model this profile talks to". Adds exactly one optional field per resource -- ChatModeProfileSettings .ConversationDeploymentName and ChatInteraction.ConversationDeploymentName -- resolved through the realtime slot, where empty means "use the site default". No transport choice is stored: whether the resolved deployment speaks natively or chains speech-to-text, chat, and text-to-speech together is read off the deployment's own CascadedRealtimeMetadata, so nothing on the profile can contradict it. ChatMode.Realtime stays gone. One shared resolution replaces the per-surface guesswork that four chat surfaces and two hubs each made for themselves, so the toggle a user sees and the session the hub will start cannot disagree. Conversation mode that resolves no realtime deployment still falls back to today's client-driven speech cascade. Migration happens at read time rather than in OnDeserialized. A resource stored before this change names its speech-to-speech model as its chat deployment and nothing else in the stored JSON says so -- the previous migration erased the RealtimeDeploymentName marker and ChatModeJsonConverter reads the old "Realtime" chat mode back as TextInput. Only the deployment's own capability distinguishes such a resource, and that needs the deployment catalog. Folding on read also keeps the stored shape honest: an editor that saves the resource persists the folded shape as a side effect of showing it, and the chat slot already excludes realtime deployments, so a typed turn lands on the site's chat default meanwhile. Because a chat interaction has no chat mode of its own -- the site holds it -- naming a conversation deployment on the interaction is itself the opt-in. Without that, an interaction would lose its voice the first time its settings were saved, since the fold only fires while the conversation deployment is empty. The client stops handing its message box to the realtime module. The module hides every control it is given when realtime takes over, and typing has to stay available during a voice conversation -- that is the whole feature. A typed message ends the live session first, because two writers into one session would interleave. Removes the hub's text-rejection guard, which existed only because a realtime chat deployment genuinely could not answer a typed turn.
…ck when it stops Follows the conversation-mode work with the editor fix, the UI the toggle deserved, and the breaking change spelled out for upgraders. The AI Template editor never showed its conversation deployment picker. The profile editor writes its chat mode options out by hand, so the value is the mode's name; the template editor generated them with Html.GetEnumSelectList, which emits the numeric value, while its own script compared against 'Conversation'. The comparison could never be true and the picker and the voice field stayed hidden. The options are written out on both template views now, and the Blazor template editor was never affected because it compares in C#. The voice toggle reads as one control instead of a labelled button crowding the row. Idle is the soundwave on its own -- the words moved to the title and a new aria-label, so the message box grew from 433px to 1042px on a 1440px screen -- and the end state keeps its words, because stopping a live conversation should be unmistakable. Send now sits with the message box it belongs to, and the voice controls follow on the right. While a session runs the message box and the send button give way to the voice settings, and the button takes the width they leave behind. Sending a typed message already ended the session first, on every surface, so the box was never a way to say something *during* one; hiding it only makes that honest. Two defects behind that, both older than this change: The session's live state was tracked twice -- the button read the realtime module's own isRealtimeActive, the controls read the host's copy -- and the WebRTC-to-WebSocket fallback moves one without the other. Whichever path repainted the button without pairing the events left a hidden message box under a button that said "Start speaking", with no way back to typing. The module now reports its state on every repaint and every host derives the controls from it, so the two cannot disagree. The module hard-coded the id of the settings popover it builds. A page that hosts both its own chat and the admin widget has two of them, so the widget's instance drove the page chat's popover. Each instance keeps the element it built. Documented in the 2.0.0 breaking changes: conversation mode resolves the realtime slot first and ResolveSlotAsync falls through to the first realtime-capable deployment, so a site with any such deployment carries conversations with it even when nothing is named. The speech-to-text plus text-to-speech cascade still serves conversation mode, but only as the fallback, and there is no per-profile switch because the transport is deliberately not stored. Installations that want voice without a speech-to-speech model should use a cascaded realtime deployment.
The guide still described the behaviour this change replaced: that a profile speaks when its chat mode is Realtime -- a mode that no longer exists -- and that the controller hides the message box because "a realtime session is audio-only". Both read as instructions to build the thing the feature removed. The overview now says what decides a spoken conversation and what carries it, a new section walks the realtime slot's chain and says plainly that any realtime-capable deployment will win it, and the user-controls section explains that typing and speaking take turns, with two screenshots of the row in each state. The controls section also tells a host how to integrate: pass only the buttons to the module and drive your own controls from onSessionStateChanged, because the activate and deactivate events are not reliably paired.
MikeAlhayek
deleted the
claude/conversation-mode-realtime-toggle-f5d10a
branch
September 21, 2026 21:06
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Design only — no feature code. Adds
docs/design/conversation-mode-realtime-toggle.md.The problem
Selecting a realtime-capable chat deployment turns the entire chat UI voice-only, across AI Chat sessions, Chat Interactions, and the chat widget. Users want typed text and realtime voice in the same thread, toggled on demand — and they want conversation mode to stop meaning "speech-to-text + text-to-speech," which is laggy enough that many avoid it.
What the doc covers
How the coupling actually works today. One field answers two questions:
AIProfile.ChatDeploymentNameis both the text model and the realtime model, read throughIsRealtimeDeploymentAsync. Voice-only is then forced in exactly three places —d-nonein the four views, theChatModesquash toTextInput, and the hubs' text-prompt rejection.What is already decoupled.
RealtimeOrchestrationRequest.RealtimeDeploymentNameexists and already resolves through the realtime slot chain (explicit name → site default → first realtime-capable), which is precisely the "default to the site's realtime model, allow an override" behavior this feature needs. The client JS already treatsrealtimeEnabledandchatModeas independent flags and never disables typing. The hubs are the only thing hardcoding the chat deployment as the realtime one.The recommendation, which differs from the original ask. Rather than storing a realtime-vs-STT/TTS transport choice plus per-branch model selections, conversation mode gets one optional field: the conversation deployment. Whether it speaks natively or is a cascade of speech-to-text + chat + text-to-speech is already a deployment property (
CascadedRealtimeMetadata), so storing it again on the profile would re-create the two-sources-of-truth problem that the realtime capability refactor deliberately removed — seeAIProfile.cs:56-90andChatModeJsonConverter.The one piece of real engineering.
DefaultRealtimeOrchestratorstarts realtime sessions with an emptyConversationHistory. That is invisible today, because a realtime profile has no text turns to miss; it becomes the feature's most obvious bug the moment both modes share a thread. Two options are weighed, with a recommendation.Also included: migration for existing realtime profiles, an 8-step build order, a verification plan, and four open questions.
Notes
CascadedRealtimeMetadatahas no admin UI today — it is seed/JSON-only. Making STT+TTS a choosable option needs a deployment editor section, scoped as independently shippable.🤖 Generated with Claude Code