Skip to content

Let one chat thread carry both typed messages and a spoken conversation - #216

Merged
MikeAlhayek merged 3 commits into
CrestApps:claude/conversation-mode-realtime-toggle-f5d10afrom
JackTelford:claude/conversation-mode-realtime-toggle-f5d10a
Sep 21, 2026
Merged

MikeAlhayek merged 3 commits into
CrestApps:claude/conversation-mode-realtime-toggle-f5d10afrom
JackTelford:claude/conversation-mode-realtime-toggle-f5d10a

Conversation

@JackTelford

Copy link
Copy Markdown
Contributor

One chat thread can now carry both typed messages and a spoken conversation, with the user deciding which to use, turn by turn. Before this, pointing a profile at a speech-to-speech model turned the whole surface voice-only: the message box was hidden, and text was rejected at the hub. Voice was something the deployment did to the UI rather than something the user chose.

What changes

Chat mode decides how a conversation is carried. The chat deployment goes back to meaning only "the text model this profile talks to". Conversation mode resolves the realtime slot and, when a realtime deployment answers, the surface offers a voice toggle next to a working message box.

One new optional field per resource, and no stored transport. ChatModeProfileSettings.ConversationDeploymentName and ChatInteraction.ConversationDeploymentName name the model that carries the conversation; left empty, the realtime slot's own chain applies. Nothing records how the conversation is carried — the repo removed a second stored answer to that question once already, and ChatMode.Realtime stays gone.

One resolver, six surfaces. ConversationModeResolutionExtensions replaced the per-surface guesswork that had four chat surfaces and two hubs each deciding separately whether a deployment meant voice. The toggle the user sees and the session the hub starts can no longer disagree.

The UI

Idle is a soundwave icon on its own, with the wording moved to the title and aria-label; on a 1440px screen that gives the message box 1042px instead of 433px. Send sits with the box it belongs to, and the voice controls follow on the right.

While a session runs, the message box and send give way to the voice settings and the button takes the width they leave. Sending a typed message already ended the session first on every surface, so the box was never a way to speak during one — hiding it only makes that honest, and it comes straight back when the session ends.

Between sessions the message box has the row, with send beside it and the voice toggle on the right:

Conversation mode idle: a wide message box, a send button, and a soundwave toggle

While a session runs, the toggle takes the row and the voice settings sit at its right-hand end:

Conversation mode speaking: a full-width End Conversation button and a settings gear, with the message box gone

Two defects this surfaced, both older than the feature

The live session was tracked twice. The button read the realtime module's isRealtimeActive; the surrounding controls read the host's own copy. The WebRTC-to-WebSocket fallback moves one without the other, so any path that repainted the button without pairing the activate/deactivate events left a hidden message box under a button reading "Start speaking" — a dead end, with no way back to typing. The module now reports its state on every repaint and every host derives the controls from it.

The settings popover had a hard-coded element id. A page hosting both its own chat and the admin widget has two, so the widget's instance drove the page chat's popover. Each instance keeps the element it built.

Also fixed: the AI Template editor never showed its conversation deployment picker. The profile editor writes its chat mode options out by hand, so the value is the mode's name; the template editor generated them with Html.GetEnumSelectList, which emits the numeric value, while its own script compared against 'Conversation'. The comparison could never be true. MVC only — the Blazor template editor compares in C#.

Breaking change

Documented in the 2.0.0 notes, and the Realtime Voice guide is rewritten to match (it still described the removed Realtime chat mode and claimed a session is audio-only). Conversation mode resolves the realtime slot first, and ResolveSlotAsync falls through to the first realtime-capable deployment — so a site with any such deployment will carry conversations with it, even with no site default configured and nothing named on the profile.

The speech-to-text plus text-to-speech cascade still serves conversation mode, but as the fallback rather than the meaning of the mode, and there is deliberately no per-profile switch back to it: the transport is not stored. To keep the cascade, have no deployment declare realtime. To speak without a speech-to-speech model, use a cascaded realtime deployment, which chains STT, chat, and TTS into one IRealtimeClient and may mix vendors.

A profile stored in the old shape is folded when it is read, not migrated on disk — an editor persists the folded shape the next time it saves, so nothing is rewritten behind the user's back. A chat interaction has no chat mode of its own, so naming a conversation deployment on it is itself the opt-in to speaking.

Verification

4003 .NET tests and 112 JS tests pass; both sample hosts build and run.

Driven in a real browser across all six surfaces — MVC AI Chat, MVC widget, MVC chat interaction, Blazor AI Chat, Blazor widget, Blazor chat interaction — checking the accessibility tree and computed geometry rather than innerText, so an aria-hidden or zero-size control would fail.

Against a live Azure realtime deployment: a session that reached "Listening", the button spanning the row, and a manual stop restoring the message box and send — then typing into it. Against live Azure Speech: a cascade conversation started, held, and stopped cleanly. Also a real completion through Azure OpenAI, and transcript survival across a reload. The legacy-shape migration and the interaction opt-in were each confirmed in the browser, and WebRTC negotiated with real microphone audio reaching the server.

Not verified: a full spoken round trip with transcribed turns interleaved in the transcript — that needs someone to actually talk to it.

Pre-existing, left alone: on the cascade path, a session that fails at start leaves the toggle reading "End Conversation" and the client streaming microphone audio indefinitely. git diff shows this change touches neither the cascade's error handling nor its stream teardown. The realtime path does not have this bug. Blazor /chat-interactions/create also still crashes its circuit on a ModelParametersEditor parameter mismatch present at the base commit.

Steps 7 (history bridging) and 8 (cascade editor) are untouched, and the resolver returns RealtimeDeploymentName where step 7 will need it.

🤖 Generated with Claude Code

A realtime-capable chat deployment turned the whole chat UI voice-only, because
AIProfile.ChatDeploymentName was simultaneously the text model and the realtime
model. Conversation mode now decides how the conversation is carried, and the
chat deployment goes back to meaning only "the text model this profile talks to".

Adds exactly one optional field per resource -- ChatModeProfileSettings
.ConversationDeploymentName and ChatInteraction.ConversationDeploymentName --
resolved through the realtime slot, where empty means "use the site default".
No transport choice is stored: whether the resolved deployment speaks natively or
chains speech-to-text, chat, and text-to-speech together is read off the
deployment's own CascadedRealtimeMetadata, so nothing on the profile can
contradict it. ChatMode.Realtime stays gone.

One shared resolution replaces the per-surface guesswork that four chat surfaces
and two hubs each made for themselves, so the toggle a user sees and the session
the hub will start cannot disagree. Conversation mode that resolves no realtime
deployment still falls back to today's client-driven speech cascade.

Migration happens at read time rather than in OnDeserialized. A resource stored
before this change names its speech-to-speech model as its chat deployment and
nothing else in the stored JSON says so -- the previous migration erased the
RealtimeDeploymentName marker and ChatModeJsonConverter reads the old "Realtime"
chat mode back as TextInput. Only the deployment's own capability distinguishes
such a resource, and that needs the deployment catalog. Folding on read also
keeps the stored shape honest: an editor that saves the resource persists the
folded shape as a side effect of showing it, and the chat slot already excludes
realtime deployments, so a typed turn lands on the site's chat default meanwhile.

Because a chat interaction has no chat mode of its own -- the site holds it --
naming a conversation deployment on the interaction is itself the opt-in.
Without that, an interaction would lose its voice the first time its settings
were saved, since the fold only fires while the conversation deployment is empty.

The client stops handing its message box to the realtime module. The module hides
every control it is given when realtime takes over, and typing has to stay
available during a voice conversation -- that is the whole feature. A typed
message ends the live session first, because two writers into one session would
interleave.

Removes the hub's text-rejection guard, which existed only because a realtime
chat deployment genuinely could not answer a typed turn.
…ck when it stops

Follows the conversation-mode work with the editor fix, the UI the toggle
deserved, and the breaking change spelled out for upgraders.

The AI Template editor never showed its conversation deployment picker. The
profile editor writes its chat mode options out by hand, so the value is the
mode's name; the template editor generated them with Html.GetEnumSelectList,
which emits the numeric value, while its own script compared against
'Conversation'. The comparison could never be true and the picker and the voice
field stayed hidden. The options are written out on both template views now, and
the Blazor template editor was never affected because it compares in C#.

The voice toggle reads as one control instead of a labelled button crowding the
row. Idle is the soundwave on its own -- the words moved to the title and a new
aria-label, so the message box grew from 433px to 1042px on a 1440px screen --
and the end state keeps its words, because stopping a live conversation should
be unmistakable. Send now sits with the message box it belongs to, and the voice
controls follow on the right.

While a session runs the message box and the send button give way to the voice
settings, and the button takes the width they leave behind. Sending a typed
message already ended the session first, on every surface, so the box was never
a way to say something *during* one; hiding it only makes that honest.

Two defects behind that, both older than this change:

The session's live state was tracked twice -- the button read the realtime
module's own isRealtimeActive, the controls read the host's copy -- and the
WebRTC-to-WebSocket fallback moves one without the other. Whichever path
repainted the button without pairing the events left a hidden message box under
a button that said "Start speaking", with no way back to typing. The module now
reports its state on every repaint and every host derives the controls from it,
so the two cannot disagree.

The module hard-coded the id of the settings popover it builds. A page that
hosts both its own chat and the admin widget has two of them, so the widget's
instance drove the page chat's popover. Each instance keeps the element it
built.

Documented in the 2.0.0 breaking changes: conversation mode resolves the
realtime slot first and ResolveSlotAsync falls through to the first
realtime-capable deployment, so a site with any such deployment carries
conversations with it even when nothing is named. The speech-to-text plus
text-to-speech cascade still serves conversation mode, but only as the fallback,
and there is no per-profile switch because the transport is deliberately not
stored. Installations that want voice without a speech-to-speech model should
use a cascaded realtime deployment.
The guide still described the behaviour this change replaced: that a profile
speaks when its chat mode is Realtime -- a mode that no longer exists -- and
that the controller hides the message box because "a realtime session is
audio-only". Both read as instructions to build the thing the feature removed.

The overview now says what decides a spoken conversation and what carries it, a
new section walks the realtime slot's chain and says plainly that any
realtime-capable deployment will win it, and the user-controls section explains
that typing and speaking take turns, with two screenshots of the row in each
state.

The controls section also tells a host how to integrate: pass only the buttons
to the module and drive your own controls from onSessionStateChanged, because
the activate and deactivate events are not reliably paired.
@MikeAlhayek
MikeAlhayek changed the base branch from main to claude/conversation-mode-realtime-toggle-f5d10a September 21, 2026 20:54
@MikeAlhayek
MikeAlhayek merged commit 521fc23 into CrestApps:claude/conversation-mode-realtime-toggle-f5d10a Sep 21, 2026
8 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants