Let the listener choose how fast the voice assistant speaks - #221
Merged
MikeAlhayek merged 2 commits intoSep 25, 2026
Merged
Conversation
Adds a Speaking speed slider (0.75x to 1.5x) to the realtime voice settings popover. Every conversation starts at 1.0x; a user's choice is saved in the browser, applied as soon as the next conversation opens, and Reset returns to 1.0x. The speed goes to the provider as a partial session.update carrying only session.audio.output.speed, so the voice is never resent. The provider applies it from the next reply and leaves the words and the voice unchanged. Cascaded deployments report that their speed cannot change, and the slider hides. Interrupted replies are now truncated on the provider's own timeline: it measures an item before the speed is applied, so the heard duration is scaled by the speed the reply was delivered at.
Contributor
Author
|
@MikeAlhayek This is ready for review Tested and working on my end. |
The realtime voice settings button showed a gear, the same glyph the sample hosts use for their own Settings pages. It now shows fa-solid fa-headset, which Font Awesome Free 7, the version the hosts load, includes. The built scripts under wwwroot are regenerated from the asset source.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A realtime conversation always spoke at the model's own pace. Nothing in the session configuration mentioned speed, so every reply played at 1.0× whether the listener wanted it quicker or needed it slower.
gpt-realtimeexposessession.audio.output.speed(0.25–1.5), which stretches the generated audio and leaves the words and the voice unchanged. This adds a Speaking speed slider to the voice settings popover, from 0.75× to 1.5×, and wires it to that setting.What the user sees
The slider sits under Assistant volume in the gear popover, which appears only during a live conversation. Every conversation starts at 1.0×. Moving the slider changes the speed from the assistant's next reply; a reply already playing keeps its speed, and the popover says so. The choice is saved with the other per-device voice preferences (
coreai.realtime.audioPrefs) and applied as soon as the next conversation opens. Reset returns to 1.0×. There is no admin setting: speed is a listener preference, like volume.How it works
Provider.
IRealtimeConversationgainsSupportsSpeechSpeedandUpdateSpeechSpeedAsync, both default interface members, so existing implementations are unaffected.DefaultRealtimeConversationsends a partialsession.updatecarrying only the speed. It follows the patternUpdateTurnDetectionAsyncalready uses, because resending the voice after the assistant has spoken is rejected:{ "type": "session.update", "session": { "type": "realtime", "audio": { "output": { "speed": 1.25 } } } }The orchestrator passes
supportsSpeechSpeed: !isCascaded. A cascaded deployment speaks through a text-to-speech leg that is never given a speed, so the slider hides instead of doing nothing. Values are clamped by the newRealtimeSpeechSpeedRange(0.75–1.5, rounded to hundredths). Azure answers anything above 1.5 withdecimal_above_max_value, which is not one of the errors the runner treats as benign, so it would reach the user.Hubs.
AIChatHubCoreandChatInteractionHubBaseboth getUpdateRealtimeSpeechSpeed(double speed), modelled onUpdateRealtimeSettings: it finds the caller's session inRealtimeSessionRegistryand applies the speed throughRealtimeSessionControl.ApplySpeechSpeedAsync. It is a separate method rather than a newUpdateRealtimeSettingsargument, because the volume slider pushes turn detection on every tick and should not resend the speed with it. TheStartRealtimeConversationandStartRealtimeWebRtcsignatures are unchanged.Session start. A new
speech_speedlifecycle event goes out just beforesession_ready. It carries the session's speed, or nothing when the speed cannot change. The client uses it to show or hide the slider and to send a saved speed straight back. For that reply to land, the session control is now published before the client hears anything:Before:
After:
Interruptions. The provider measures an assistant item before the speed is applied, so
conversation.item.truncatehas to name the un-sped position. The runner records the speed each reply started at and scales what was heard:Before:
After:
Without the scaling, a barge-in at 1.5× would cut a third of what the user actually heard out of the model's context. At 0.75× the cut would land past the end of the item and be refused with an error the runner already filters as benign, so truncation would quietly stop working.
Docs.
realtime-voice.mdgains the new user control, thespeech_speedevent and the truncation rule.Verified against Azure
gpt-realtimeBefore building this, I measured it directly against a
gpt-realtimedeployment reading the same 24-word sentence:session.updatedechoesaudio.output.speed, and the transcript is word-for-word identical at every speed.speed: 2.0is rejected withExpected a value <= 1.5.Audio content of 9050ms, and 13,568 ms delivered at 0.75× as10250ms.Testing
npm run test:unit: 112 passed. Realtime client browser tests: 18/18 passed in Chromium; the Firefox project could not run because Firefox is not installed on this machine.gpt-realtime, ran on an earlier revision of this branch that also had per-profile defaults (since removed). Every slider move arrived asUpdateRealtimeSpeechSpeed, a saved speed was sent back as the next conversation opened, and nothing was logged as a warning or an error.Why users will like it
Everyone gets to hear the assistant at the pace that suits them: faster for someone who knows the material and wants to get through it, slower for someone following along in a second language, taking notes, or who finds fast speech hard to keep up with. It is one control, in the place the voice settings already live, remembered per browser. Nothing changes for anyone who never touches it.
🤖 Generated with Claude Code