Skip to content

Design: typed text and realtime voice on the same chat UI - #214

Merged
MikeAlhayek merged 5 commits into
mainfrom
claude/conversation-mode-realtime-toggle-f5d10a
Sep 21, 2026
Merged

MikeAlhayek merged 5 commits into
mainfrom
claude/conversation-mode-realtime-toggle-f5d10a

Conversation

@MikeAlhayek

Copy link
Copy Markdown
Member

Design only — no feature code. Adds docs/design/conversation-mode-realtime-toggle.md.

The problem

Selecting a realtime-capable chat deployment turns the entire chat UI voice-only, across AI Chat sessions, Chat Interactions, and the chat widget. Users want typed text and realtime voice in the same thread, toggled on demand — and they want conversation mode to stop meaning "speech-to-text + text-to-speech," which is laggy enough that many avoid it.

What the doc covers

How the coupling actually works today. One field answers two questions: AIProfile.ChatDeploymentName is both the text model and the realtime model, read through IsRealtimeDeploymentAsync. Voice-only is then forced in exactly three places — d-none in the four views, the ChatMode squash to TextInput, and the hubs' text-prompt rejection.

What is already decoupled. RealtimeOrchestrationRequest.RealtimeDeploymentName exists and already resolves through the realtime slot chain (explicit name → site default → first realtime-capable), which is precisely the "default to the site's realtime model, allow an override" behavior this feature needs. The client JS already treats realtimeEnabled and chatMode as independent flags and never disables typing. The hubs are the only thing hardcoding the chat deployment as the realtime one.

The recommendation, which differs from the original ask. Rather than storing a realtime-vs-STT/TTS transport choice plus per-branch model selections, conversation mode gets one optional field: the conversation deployment. Whether it speaks natively or is a cascade of speech-to-text + chat + text-to-speech is already a deployment property (CascadedRealtimeMetadata), so storing it again on the profile would re-create the two-sources-of-truth problem that the realtime capability refactor deliberately removed — see AIProfile.cs:56-90 and ChatModeJsonConverter.

The one piece of real engineering. DefaultRealtimeOrchestrator starts realtime sessions with an empty ConversationHistory. That is invisible today, because a realtime profile has no text turns to miss; it becomes the feature's most obvious bug the moment both modes share a thread. Two options are weighed, with a recommendation.

Also included: migration for existing realtime profiles, an 8-step build order, a verification plan, and four open questions.

Notes

  • CascadedRealtimeMetadata has no admin UI today — it is seed/JSON-only. Making STT+TTS a choosable option needs a deployment editor section, scoped as independently shippable.
  • Today's client-driven STT/TTS conversation mode is deliberately kept, because the cascade requires a streaming transcription client while the current path works with any speech-to-text deployment.

🤖 Generated with Claude Code

MikeAlhayek and others added 5 commits September 21, 2026 09:38
Today a realtime-capable chat deployment turns the whole chat UI voice-only,
because one field answers two questions: AIProfile.ChatDeploymentName is both
the text model and the realtime model. Users want to type and speak in the same
thread, toggled on demand.

This is design only. It works through what the coupling actually is (three
places: the views' d-none, the ChatMode squash, and the hubs' text rejection),
notes that the orchestration layer is already decoupled via
RealtimeOrchestrationRequest.RealtimeDeploymentName and the realtime slot chain,
and recommends a single new optional field -- the conversation deployment --
rather than storing a realtime-vs-STT/TTS transport choice that
CascadedRealtimeMetadata already answers at the deployment level.

Also records the one piece of real engineering: DefaultRealtimeOrchestrator
starts realtime sessions with an empty ConversationHistory, which is invisible
today but becomes the feature's most obvious bug once both modes share a thread.
A realtime-capable chat deployment turned the whole chat UI voice-only, because
AIProfile.ChatDeploymentName was simultaneously the text model and the realtime
model. Conversation mode now decides how the conversation is carried, and the
chat deployment goes back to meaning only "the text model this profile talks to".

Adds exactly one optional field per resource -- ChatModeProfileSettings
.ConversationDeploymentName and ChatInteraction.ConversationDeploymentName --
resolved through the realtime slot, where empty means "use the site default".
No transport choice is stored: whether the resolved deployment speaks natively or
chains speech-to-text, chat, and text-to-speech together is read off the
deployment's own CascadedRealtimeMetadata, so nothing on the profile can
contradict it. ChatMode.Realtime stays gone.

One shared resolution replaces the per-surface guesswork that four chat surfaces
and two hubs each made for themselves, so the toggle a user sees and the session
the hub will start cannot disagree. Conversation mode that resolves no realtime
deployment still falls back to today's client-driven speech cascade.

Migration happens at read time rather than in OnDeserialized. A resource stored
before this change names its speech-to-speech model as its chat deployment and
nothing else in the stored JSON says so -- the previous migration erased the
RealtimeDeploymentName marker and ChatModeJsonConverter reads the old "Realtime"
chat mode back as TextInput. Only the deployment's own capability distinguishes
such a resource, and that needs the deployment catalog. Folding on read also
keeps the stored shape honest: an editor that saves the resource persists the
folded shape as a side effect of showing it, and the chat slot already excludes
realtime deployments, so a typed turn lands on the site's chat default meanwhile.

Because a chat interaction has no chat mode of its own -- the site holds it --
naming a conversation deployment on the interaction is itself the opt-in.
Without that, an interaction would lose its voice the first time its settings
were saved, since the fold only fires while the conversation deployment is empty.

The client stops handing its message box to the realtime module. The module hides
every control it is given when realtime takes over, and typing has to stay
available during a voice conversation -- that is the whole feature. A typed
message ends the live session first, because two writers into one session would
interleave.

Removes the hub's text-rejection guard, which existed only because a realtime
chat deployment genuinely could not answer a typed turn.
…ck when it stops

Follows the conversation-mode work with the editor fix, the UI the toggle
deserved, and the breaking change spelled out for upgraders.

The AI Template editor never showed its conversation deployment picker. The
profile editor writes its chat mode options out by hand, so the value is the
mode's name; the template editor generated them with Html.GetEnumSelectList,
which emits the numeric value, while its own script compared against
'Conversation'. The comparison could never be true and the picker and the voice
field stayed hidden. The options are written out on both template views now, and
the Blazor template editor was never affected because it compares in C#.

The voice toggle reads as one control instead of a labelled button crowding the
row. Idle is the soundwave on its own -- the words moved to the title and a new
aria-label, so the message box grew from 433px to 1042px on a 1440px screen --
and the end state keeps its words, because stopping a live conversation should
be unmistakable. Send now sits with the message box it belongs to, and the voice
controls follow on the right.

While a session runs the message box and the send button give way to the voice
settings, and the button takes the width they leave behind. Sending a typed
message already ended the session first, on every surface, so the box was never
a way to say something *during* one; hiding it only makes that honest.

Two defects behind that, both older than this change:

The session's live state was tracked twice -- the button read the realtime
module's own isRealtimeActive, the controls read the host's copy -- and the
WebRTC-to-WebSocket fallback moves one without the other. Whichever path
repainted the button without pairing the events left a hidden message box under
a button that said "Start speaking", with no way back to typing. The module now
reports its state on every repaint and every host derives the controls from it,
so the two cannot disagree.

The module hard-coded the id of the settings popover it builds. A page that
hosts both its own chat and the admin widget has two of them, so the widget's
instance drove the page chat's popover. Each instance keeps the element it
built.

Documented in the 2.0.0 breaking changes: conversation mode resolves the
realtime slot first and ResolveSlotAsync falls through to the first
realtime-capable deployment, so a site with any such deployment carries
conversations with it even when nothing is named. The speech-to-text plus
text-to-speech cascade still serves conversation mode, but only as the fallback,
and there is no per-profile switch because the transport is deliberately not
stored. Installations that want voice without a speech-to-speech model should
use a cascaded realtime deployment.
The guide still described the behaviour this change replaced: that a profile
speaks when its chat mode is Realtime -- a mode that no longer exists -- and
that the controller hides the message box because "a realtime session is
audio-only". Both read as instructions to build the thing the feature removed.

The overview now says what decides a spoken conversation and what carries it, a
new section walks the realtime slot's chain and says plainly that any
realtime-capable deployment will win it, and the user-controls section explains
that typing and speaking take turns, with two screenshots of the row in each
state.

The controls section also tells a host how to integrate: pass only the buttons
to the module and drive your own controls from onSessionStateChanged, because
the activate and deactivate events are not reliably paired.
@MikeAlhayek
MikeAlhayek merged commit ad0a796 into main Sep 21, 2026
10 checks passed
@MikeAlhayek
MikeAlhayek deleted the claude/conversation-mode-realtime-toggle-f5d10a branch September 21, 2026 21:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants