Skip to content

Let the listener choose how fast the voice assistant speaks - #221

Merged
MikeAlhayek merged 2 commits into
CrestApps:mainfrom
JackTelford:jt-add-realtime-speed-speech
Sep 25, 2026
Merged

MikeAlhayek merged 2 commits into
CrestApps:mainfrom
JackTelford:jt-add-realtime-speed-speech

Conversation

@JackTelford

Copy link
Copy Markdown
Contributor

A realtime conversation always spoke at the model's own pace. Nothing in the session configuration mentioned speed, so every reply played at 1.0× whether the listener wanted it quicker or needed it slower. gpt-realtime exposes session.audio.output.speed (0.25–1.5), which stretches the generated audio and leaves the words and the voice unchanged. This adds a Speaking speed slider to the voice settings popover, from 0.75× to 1.5×, and wires it to that setting.

What the user sees

The slider sits under Assistant volume in the gear popover, which appears only during a live conversation. Every conversation starts at 1.0×. Moving the slider changes the speed from the assistant's next reply; a reply already playing keeps its speed, and the popover says so. The choice is saved with the other per-device voice preferences (coreai.realtime.audioPrefs) and applied as soon as the next conversation opens. Reset returns to 1.0×. There is no admin setting: speed is a listener preference, like volume.

How it works

Provider. IRealtimeConversation gains SupportsSpeechSpeed and UpdateSpeechSpeedAsync, both default interface members, so existing implementations are unaffected. DefaultRealtimeConversation sends a partial session.update carrying only the speed. It follows the pattern UpdateTurnDetectionAsync already uses, because resending the voice after the assistant has spoken is rejected:

{ "type": "session.update", "session": { "type": "realtime", "audio": { "output": { "speed": 1.25 } } } }

The orchestrator passes supportsSpeechSpeed: !isCascaded. A cascaded deployment speaks through a text-to-speech leg that is never given a speed, so the slider hides instead of doing nothing. Values are clamped by the new RealtimeSpeechSpeedRange (0.75–1.5, rounded to hundredths). Azure answers anything above 1.5 with decimal_above_max_value, which is not one of the errors the runner treats as benign, so it would reach the user.

Hubs. AIChatHubCore and ChatInteractionHubBase both get UpdateRealtimeSpeechSpeed(double speed), modelled on UpdateRealtimeSettings: it finds the caller's session in RealtimeSessionRegistry and applies the speed through RealtimeSessionControl.ApplySpeechSpeedAsync. It is a separate method rather than a new UpdateRealtimeSettings argument, because the volume slider pushes turn detection on every tick and should not resend the speed with it. The StartRealtimeConversation and StartRealtimeWebRtc signatures are unchanged.

Session start. A new speech_speed lifecycle event goes out just before session_ready. It carries the session's speed, or nothing when the speed cannot change. The client uses it to show or hide the slider and to send a saved speed straight back. For that reply to land, the session control is now published before the client hears anything:

Before:

await sink.SessionReadyAsync(context.SessionId, cancellationToken);
context.OnSessionStarted?.Invoke(new RealtimeSessionControl(conversation, context));

After:

context.OnSessionStarted?.Invoke(new RealtimeSessionControl(conversation, context));
await sink.SpeechSpeedAsync(context.SessionId, conversation.SupportsSpeechSpeed ? context.SpeechSpeed : null, cancellationToken);
await sink.SessionReadyAsync(context.SessionId, cancellationToken);

Interruptions. The provider measures an assistant item before the speed is applied, so conversation.item.truncate has to name the un-sped position. The runner records the speed each reply started at and scales what was heard:

Before:

var heardMs = playback.SentMs - sink.PendingPlaybackMs - ClientJitterBufferMs;

After:

var heardMs = (int)((playback.SentMs - sink.PendingPlaybackMs - ClientJitterBufferMs) * playback.Speed);

Without the scaling, a barge-in at 1.5× would cut a third of what the user actually heard out of the model's context. At 0.75× the cut would land past the end of the item and be refused with an error the runner already filters as benign, so truncation would quietly stop working.

Docs. realtime-voice.md gains the new user control, the speech_speed event and the truncation rule.

Verified against Azure gpt-realtime

Before building this, I measured it directly against a gpt-realtime deployment reading the same 24-word sentence:

Speed Audio length vs 1.0×
1.0× 8.90 s —
1.5× 6.02 s 1.48× faster
0.75× 11.64 s 0.76×
  • session.updated echoes audio.output.speed, and the transcript is word-for-word identical at every speed.
  • A speed sent while a reply is generating is accepted without error and applies to the next reply.
  • speed: 2.0 is rejected with Expected a value <= 1.5.
  • The truncation timeline is the un-sped one. 6,020 ms delivered at 1.5× is reported by the server as Audio content of 9050ms, and 13,568 ms delivered at 0.75× as 10250ms.

Testing

  • The solution builds with 0 warnings and 0 errors.
  • Existing realtime tests: 223 passed. npm run test:unit: 112 passed. Realtime client browser tests: 18/18 passed in Chromium; the Firefox project could not run because Firefox is not installed on this machine.
  • A live session in the MVC sample host, over WebRTC against Azure gpt-realtime, ran on an earlier revision of this branch that also had per-profile defaults (since removed). Every slider move arrived as UpdateRealtimeSpeechSpeed, a saved speed was sent back as the next conversation opened, and nothing was logged as a warning or an error.
  • No new tests are included.

Why users will like it

Everyone gets to hear the assistant at the pace that suits them: faster for someone who knows the material and wants to get through it, slower for someone following along in a second language, taking notes, or who finds fast speech hard to keep up with. It is one control, in the place the voice settings already live, remembered per browser. Nothing changes for anyone who never touches it.

🤖 Generated with Claude Code

Adds a Speaking speed slider (0.75x to 1.5x) to the realtime voice settings
popover. Every conversation starts at 1.0x; a user's choice is saved in the
browser, applied as soon as the next conversation opens, and Reset returns to
1.0x.

The speed goes to the provider as a partial session.update carrying only
session.audio.output.speed, so the voice is never resent. The provider applies
it from the next reply and leaves the words and the voice unchanged. Cascaded
deployments report that their speed cannot change, and the slider hides.

Interrupted replies are now truncated on the provider's own timeline: it
measures an item before the speed is applied, so the heard duration is scaled
by the speed the reply was delivered at.
@JackTelford

Copy link
Copy Markdown
Contributor Author

@MikeAlhayek This is ready for review Tested and working on my end.

The realtime voice settings button showed a gear, the same glyph the sample
hosts use for their own Settings pages. It now shows fa-solid fa-headset,
which Font Awesome Free 7, the version the hosts load, includes. The built
scripts under wwwroot are regenerated from the asset source.
@MikeAlhayek
MikeAlhayek merged commit 7296404 into CrestApps:main Sep 25, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants