Skip to content

FE-1580: Reconcile Voice turn behavior on the shared Brunch conversation - #9564

Merged
lunelson merged 14 commits into
mainfrom
ln/fe-1580-reconcile-voice-resumable-workpiece
Sep 8, 2026
Merged

FE-1580: Reconcile Voice turn behavior on the shared Brunch conversation#9564
lunelson merged 14 commits into
mainfrom
ln/fe-1580-reconcile-voice-resumable-workpiece

Conversation

@lunelson

@lunelson lunelson commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Important

Current status: Merged into main as fb96f213188da885becd3248fdbe6e84abb65877. Focused follow-up #9588 ports the settlement behavior from #9531 commit 9415e1b007 that this reconciliation did not import. Treat #9588 acceptance as the remaining settlement gate before recutting #9538, #9550, and #9571 from the latest main.

🌟 What is the purpose of this PR?

Make Voice's completed-transcript, half-duplex experience work on the same canonical Brunch conversation that Mission 6 already uses for a resumable workpiece and browser mutation. In the local Petrinaut Brunch panel you can type, then speak, let Brunch change the prepared net, use Your turn or Stop, and reopen the same conversation without replaying speech or duplicating the change.

This imports KA's Voice hardening onto Mission 6's repaired fixture path and reconciles the joins that neither parent owned: deferred browser-tool execution, continuation ownership, Stop withholding, canonical stopped-entry projection, and surviving Voice origins after history fold.

What the accepted proof establishes: scoped automated suites plus the 2026-09-07 owner witness cover half-duplex admission, a causally necessary spoken browser mutation, causal per-step client results, Your turn, coherent Tab-B resume, active-submission durable Stop, Tab-C stopped-entry recovery, playback controls, and no duplicate/autoplay. It does not claim direct spoken-user attribution after hydration, durable recovery of locally withheld work after a settled tool-call step, comparative latency, or remote deployment.

🔗 Related links

🚫 Blocked by

  • The parent resumable-workpiece PR, #9537
  • Mission 6b's narrowed local claim was accepted by Lu on 2026-09-07
  • KA's original PR remains untouched; retirement requires separate authorization

🔍 What does this change?

  • Imports KA's Voice contribution onto the Mission 6 branch with attribution, then reconciles it with Mission 6's fixture, catalogue, and coherent-bundle path instead of replacing either.
  • Keeps the panel's useChat / Flue browser transport as the only Voice admission door; Realtime has no tools and cannot submit.
  • Holds the composer busy across deferred browser-tool execution and automatic continuation; Stop withholds work that has not been admitted; Your turn cancels audio without aborting admitted Brunch work.
  • Projects canonically aborted entries as stopped history, skips runnable parts of those messages on reopen, and preserves folded Voice origins on continuations.
  • Surfaces automatic-tool failures to Voice, keeps compact/expanded Voice setup and exact full-response / marked-question replay, and adds a Petrinaut changeset for the consumer-visible Voice/stop behavior.
🏗️ Agent notes

Mission authority is libs/@hashintel/brunch-agent/MISSION.md (accepted with explicit limitations on 2026-09-07). This is the explicit exception to one new issue per mission: the owner directed a PR that references FE-1580 without rewriting that issue.

Imperative. Make KA's completed-transcript, half-duplex Voice experience work safely over Mission 6's resumable browser mutations and coherent workpiece/document recovery. Preserve both capabilities. Distinguish committed prose, submission settlement, pending browser work, coherent document settlement, and terminal provider output at the shared boundaries.

Throughline. Completed current-turn microphone transcript → shared panel submitVoiceInputWithAdmission / useChat → browser ChatTransport over the memoized Flue client → same-origin /agents/chat/:instanceId and mounted Brunch ChatAgent → committed canonical prose, hidden server question marker, browser-tool requests → existing canonical browser validation and effects → original call-id outputs resume the same conversation → canonical speech queue and acknowledged cancellation → coherent workpiece/document settlement → canonical history reopen and another real turn.

Proof. The accepted owner witness used the real microphone and prepared fixture. An SDCPN negative control performed one read and no mutation; an explicit spoken confirmation then produced one addArc, a separate verification read, two model workpiece revisions, coherent revision 2, one audible reply, Tab-B continuation, canonical durable abort and Tab-C stopped-entry recovery. The witness first exposed cumulative cross-step client results and a non-causal prepared answer; both received failing-first regressions and repairs. Direct spoken-user hydration attribution, post-settlement durable withholding and comparative latency were explicitly deferred with no claim.

Constraints. One conversation/log and the mounted Flue route; no direct Voice send or live brunch_ask. Realtime has tool_choice: none and create_response: false. Only durably completed, submission-correlated canonical assistant prose may speak. Local playback cancellation, local withheld browser work, and durable Flue abortion stay distinct. Preserve Mission 6 fixture identity, scoped catalogue, no-op honesty, and prior-coherent-bundle refusal. KA's original branch/PR stay untouched.

Fog-line. Question-marker compliance remains a model limitation. Direct-user attribution after hydration and comparative latency are explicitly deferred. Stop is durable for active Flue submissions; locally withheld browser work after a settled tool-call step may reappear as pending after reopen. The negative-control answer was too verbose and durable Stop was poorly discoverable while Voice remained active.

Stop or reorient. Stop if another conversation route appears, admission auto-retries, speech is rewritten, Stop lets withheld work execute, failures disappear, or coherent settlement is falsely reported. Reorient rather than invent a durable marker, forge aborted settlements, or disable ordinary pending-tool recovery to paper over the local-Stop/reopen gap.

Deferred. Mission 7 consumes this local reconciliation and must re-pin the prompt/tool baseline before paid runs; it cannot inherit acceptance. Retirement of KA's original PR needs separate authorization.

Implementation record

  • Source contribution range: 58f7584080..be56a18ff0, imported with attribution; see import.md.
  • Latest restack verification is in docs/evidence/implementations/voice-resumable-reconciliation/main-restack.md.
  • Newly confirmed limitation: after a completed Flue tool-call step, parent Stop can return already-settled. The panel can withhold pending browser work locally, but a fresh process cannot infer that withholding from the snapshot, so reopen can recover the tool as pending work. User docs warn about this. It is an owner reorientation point, not a solved projection case.
  • Direct spoken-user Voice chip after snapshot-only reopen is still unsupported by SDK 2.0.3.

Pre-Merge Checklist 🚀

🚢 Has this modified a publishable library?

This PR:

  • modifies an npm-publishable library and I have added a changeset file(s) (@hashintel/petrinaut: .changeset/flue-voice-safety.md; the @hashintel/brunch-agent* packages are private)

📜 Does this require a change to the docs?

The changes in this PR:

  • require changes to docs which are made as part of this PR (libs/@hashintel/petrinaut/docs/ai-assistant.md, apps/petrinaut-website/README.md, and the Brunch reconciliation evidence)

🕸️ Does this require a change to the Turbo Graph?

The changes in this PR:

  • do not affect the execution graph (dev:brunch now also builds @apps/brunch-agent first; no turbo.json change)

⚠️ Known issues

  • Local Stop is not a durable stopped record. If Flue has already completed a tool-call step, Stop can withhold the browser continuation in this process and still leave that pending tool recoverable after reopen. Canonically aborted submissions project metadata.stopped and are skipped; this local-withholding case is not the same thing.
  • Direct spoken-user attribution after snapshot reopen is unsupported on the current AI SDK. The chip cannot be invented locally.
  • Direct spoken-user attribution after hydration is unsupported. Canonical text survives; the live VOICE chips disappear after fresh hydration.
  • Post-settlement local withholding is not durable. Stop is durable for active Flue submissions; locally withheld pending browser work may reappear after reopen.
  • No comparative latency claim. The 10 donor + 10 candidate campaign was explicitly deferred.
  • Interaction strain. The negative-control answer was too verbose, and durable Stop required exiting Voice mode before the streaming Stop action was discoverable.

🐾 Next steps

  • Mission 7 (#9562) is restacked on this accepted narrowed foundation; its own scenario evidence remains required.
  • Re-enter the three deferred claims only under the conditions recorded in the owner witness.
  • Separate authorization is still required to retire KA's original PR.

🛡 What tests cover this?

  • Transport: deterministic client-tool keys after reordering, surviving folded Voice origins, aborted-entry projection.
  • Petrinaut panel: busy-across-continuation, Stop-before-deferred-execution, textless automatic-tool failure, StrictMode continuation, conversation-switch isolation.
  • Website Voice: half-duplex handoff, completed-transcript admission, canonical speech selection, cancellation acknowledgement, combined voice-browser-tools.integration.test.tsx.
  • Core: question-marker helper used by exact-question replay.
  • Scoped verification recorded 39 uncached build/test/type/lint tasks and 1,318 tests; full GitHub CI and live audio evidence were not run for that record.

❓ How to test this?

  1. Run yarn dev:brunch.
  2. Open http://127.0.0.1:4915/?brunch-fixture=crew-reservation-v1.
  3. Make a typed turn, then a spoken confirmation; watch the single crew-reservation arc and coherent bundle settle.
  4. During another response, use Your turn, wait for safe fresh capture, and speak again.
  5. Separately Stop before completion.
  6. Reopen the same URL in a second tab; confirm conversation and net, then continue without duplicate preparation, mutation, or autoplay.
  7. Inspect compact/expanded Voice, exact full-response and marked-question replay, and a visible tool failure.

This local demo and its acceptance gates, not merely green unit tests, define the visible advance.

📹 Demo

The accepted owner witness records the real microphone/browser path, canonical summary, screenshots, explicit deferrals, and evidence-bundle limitation. The complete pre-registered telemetry/network/latency bundle was not retained and is not inferred.

@vercel

vercel Bot commented Sep 7, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
hash Ready Ready Preview Sep 8, 2026 3:09pm UTC
petrinaut Ready Ready Preview Sep 8, 2026 3:09pm UTC
petrinaut-docs Ready Ready Preview Sep 8, 2026 3:09pm UTC
1 Skipped Deployment
Project Deployment Actions Updated
hashdotdesign-tokens Ignored Ignored Preview Sep 8, 2026 3:09pm UTC

Request Review

@github-actions github-actions Bot added area/infra Relates to version control, CI, CD or IaC (area) area/libs Relates to first-party libraries/crates/packages (area) type/eng > frontend Owned by the @frontend team area/tests New or updated tests area/apps area/apps > hash.design Affects the `hash.design` design site (app) labels Sep 7, 2026
@lunelson
lunelson marked this pull request as ready for review September 7, 2026 12:57
Copilot AI balanced review requested due to automatic review settings September 7, 2026 12:57
@cursor

cursor Bot commented Sep 7, 2026

Copy link
Copy Markdown

PR Summary

High Risk
Changes half-duplex Realtime admission, Flue Stop/admission correlation, and canonical history projection for Voice—mistakes could duplicate turns, speak wrong text, or submit audio during assistant playback.

Overview
Reconciles Petrinaut Voice with Mission 6’s canonical Brunch/Flue path: spoken answers still go through the panel transport only, while Realtime becomes a half-duplex media layer (no tools, transcription-only admission) with stricter input/output gating and async playback cancellation.

Brunch question marking replaces interactive brunch_ask: core exposes brunch_mark_question plus a hidden data-brunch-question marker; the transport and history projection hide the tool in the UI but keep the marker for Repeat question replay when it matches finalized prose. Canonical speech selection splits full-response text from the marked question segment.

Panel ↔ Voice wiring expands BrunchPanelConversationTracker (admission failures, response start/complete, stop requested) and passes those hooks into VoiceInterviewControl. Durable Flue Stop is separated from local audio cancel; dev:brunch now builds @apps/brunch-agent and local Vite merge preserves website API plugins for Voice.

Docs/changeset describe busy composer across browser-tool continuations, stopped history, surviving Voice client-tool origins, and known limits (e.g. post-settlement Stop withholding, spoken-user chip after hydration). Tests and fixture copy tighten question-marker, realtime safety, and prepared-crew claim boundaries.

Reviewed by Cursor Bugbot for commit 743c3c8. Bugbot is set up for automated code reviews on this repo. Configure here.

@codecov

codecov Bot commented Sep 7, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 64.51%. Comparing base (66aec22) to head (de65620).

Additional details and impacted files
@@                             Coverage Diff                              @@
##           ln/fe-1575-resumable-workpiece-petrinaut    #9564      +/-   ##
============================================================================
- Coverage                                     65.93%   64.51%   -1.43%     
============================================================================
  Files                                          1886     1216     -670     
  Lines                                        198341   110645   -87696     
  Branches                                       8236     5890    -2346     
============================================================================
- Hits                                         130774    71381   -59393     
+ Misses                                        66037    38317   -27720     
+ Partials                                       1530      947     -583     
Flag Coverage Δ
apps.hash-ai-worker-py ?
apps.hash-ai-worker-ts 1.99% <ø> (ø)
apps.hash-api 15.57% <ø> (ø)
apps.hash-graph 12.54% <ø> (ø)
backend-integration-tests ?
blockprotocol.type-system 38.15% <ø> (ø)
deer ?
error-stack ?
local.claude-hooks 0.00% <ø> (ø)
local.harpc-client 51.49% <ø> (ø)
local.hash-backend-utils 3.27% <ø> (ø)
local.hash-graph-sdk 10.02% <ø> (ø)
local.hash-isomorphic-utils 12.22% <ø> (ø)
local.hash-subgraph ?
rust.antsi 2.36% <ø> (ø)
rust.deer ?
rust.error-stack 90.81% <ø> (ø)
rust.harpc-codec 84.70% <ø> (ø)
rust.harpc-net 96.19% <ø> (-0.05%) ⬇️
rust.harpc-tower 67.03% <ø> (ø)
rust.harpc-types 0.00% <ø> (ø)
rust.harpc-wire-protocol 92.23% <ø> (ø)
rust.hash-codec ?
rust.hash-config 81.14% <ø> (ø)
rust.hash-graph-api 19.71% <ø> (ø)
rust.hash-graph-atlas ?
rust.hash-graph-authentication ?
rust.hash-graph-authorization 63.14% <ø> (ø)
rust.hash-graph-embeddings 91.88% <ø> (ø)
rust.hash-graph-postgres-store ?
rust.hash-graph-store 48.41% <ø> (ø)
rust.hash-graph-temporal-versioning 50.18% <ø> (ø)
rust.hash-graph-types 0.00% <ø> (ø)
rust.hash-graph-validation 84.71% <ø> (ø)
rust.hash-middleware 90.92% <ø> (ø)
rust.hashql-ast ?
rust.hashql-compiletest 28.39% <ø> (ø)
rust.hashql-core 78.92% <ø> (ø)
rust.hashql-diagnostics 72.51% <ø> (ø)
rust.hashql-eval 79.82% <ø> (ø)
rust.hashql-hir 89.09% <ø> (ø)
rust.hashql-mir 87.92% <ø> (ø)
rust.hashql-syntax-jexpr 94.04% <ø> (ø)
rust.sarif ?
sarif ?
tests.hash-backend-integration ?
unit-tests ?

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@lunelson lunelson changed the title Cut Mission 6b Voice reconciliation authority FE-1580: Reconcile Voice turn behavior on the shared Brunch conversation Sep 7, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It spans transport durability, Voice state, browser tools, and public UI behavior while required human microphone, reload, and latency gates remain open.

Pull request overview

Reconciles Petrinaut Voice with the canonical, resumable Brunch/Flue conversation path.

Changes:

  • Adds half-duplex Voice handoff, replay controls, compact docks, and transcript authority.
  • Strengthens transport idempotency, error handling, history projection, and stopped/Voice metadata.
  • Adds question markers, persistent diagnostics, extensive tests, documentation, and mission evidence.
File summaries
File Description
package.json Prebuilds Brunch dependencies for local development.
.../ai-assistant-panel/types.ts Extends message metadata for Voice origins and stopped responses.
.../voice-dock/transcription-icon.tsx Removes the obsolete transcription icon.
.../voice-dock/playback-menu.tsx Adds Voice replay actions.
.../ai-assistant-contents/voice-dock.tsx Adds collapse, replay, handoff, and notice controls.
.../ai-assistant-contents/tool-list.tsx Displays complete tool errors inline.
.../defer-voice-messages.ts Removes deferred transcript behavior.
.../defer-voice-messages.test.ts Removes obsolete deferral tests.
.../ai-assistant-contents.stories.tsx Adds compact and collapsed Voice stories.
.../components/voice-session-labels.ts Adds labels for new Voice controls.
.../types/ai-assistant-composer-control.ts Extends public Voice controls and stopped state.
.../voice-session/use-voice-session.ts Adds granular Voice capability hooks.
.../voice-session/types.ts Extends Voice session state.
.../voice-session/store.ts Extends Voice actions.
.../notifications/toaster.tsx Adds persistent, detailed, copyable notifications.
.../notifications/provider.tsx Supplies notification details and persistent errors.
.../notifications/provider.test.tsx Tests notification duration and detail behavior.
.../notifications/context.ts Adds notification detail support.
.../panda-preset.ts Removes an unused Voice animation.
.../docs/ai-assistant.md Documents the revised Voice and error UX.
.../transport-aisdk/test/ui-stream.test.ts Tests hidden tools and error projection.
.../transport-aisdk/test/transcript.test.ts Tests durable Voice and stopped metadata.
.../transport-aisdk/test/chat-transport.test.ts Tests admission identity, failures, and callbacks.
.../transport-aisdk/src/ui-stream.ts Hides internal tools and preserves error details.
.../transport-aisdk/src/transcript.ts Reconstructs history metadata and hides marker tools.
.../transport-aisdk/src/index.ts Adds typed admission failures and deterministic delivery.
.../transport-aisdk/src/error-text.ts Adds bounded error serialization.
.../core/vite.config.ts Builds the question-marker entry point.
.../core/test/question-marker.test.ts Tests marker schemas, tool behavior, and prompt instructions.
.../core/src/question-marker.ts Defines the durable question-marker contract.
.../core/src/prompts/SYSTEM.md Instructs Brunch to mark direct questions.
.../core/src/index.ts Exports question-marker APIs.
.../core/src/flue.ts Mounts the question-marker tool.
.../core/package.json Exposes the question-marker subpath.
.../MISSION.next.md Updates the mission dependency and planning record.
.../voice-resumable-reconciliation/verification.md Records verification results and remaining gates.
.../voice-resumable-reconciliation/main-restack.md Records the restack and route verification.
.../voice-resumable-reconciliation/import.md Records imported source provenance.
.../mission-5-voice-safety-parity/witness-blocker.md Documents the outstanding human witness.
.../mission-5-voice-safety-parity/provenance-blocker.md Documents the direct-user provenance limitation.
.../mission-5-voice-safety-parity/donor-behavior-matrix.md Records adopted Voice behavior and acceptance status.
.../mission-5-question-marker-and-provenance-decision-2026-09-04.md Records question-marker and provenance decisions.
.../server/voice/openai-voice-policy.ts Makes Realtime transcription-only and half-duplex.
.../server/voice/openai-voice-policy.test.ts Tests the revised Realtime policy.
.../server/voice/openai-realtime-call.test.ts Tests forwarding the half-duplex policy.
.../voice-interview/voice-session-state.ts Projects replay, handoff, and notice state.
.../voice-interview/voice-session-state.test.ts Tests the new projected state.
.../voice-interview/voice-interview-control.test.tsx Tests admission, setup, microphone, and controls.
.../voice-interview/voice-browser-tools.integration.test.tsx Exercises the combined Voice/browser-tool lifecycle.
.../voice-interview/canonical-speech.ts Selects exact canonical response and question speech.
.../voice-interview/canonical-speech.test.ts Tests exact marker-based question selection.
.../local-storage-demo/use-flue-chat-history.ts Applies hidden-tool and history projection rules.
.../local-storage-demo/use-flue-chat-history.test.ts Tests persisted Voice origins across reopen.
.../local-storage-demo/local-storage-demo-app.tsx Wires Voice lifecycle tracking and removes brunch_ask.
.../local-storage-demo/local-storage-demo-app.test.tsx Tests production Voice registration and durable Stop.
.../local-storage-demo/brunch-panel-transport.ts Tracks response lifecycle and admission failures.
.../local-storage-demo/brunch-panel-transport.test.ts Tests tracker events and failure propagation.
apps/petrinaut-website/README.md Documents current Voice behavior and limitations.
apps/brunch-agent/test/petrinaut-chat.test.ts Verifies marker persistence and hiding.
apps/brunch-agent/test/petrinaut-chat.integration.ts Extends the real Flue integration scenario.
apps/brunch-agent/test/petrinaut-chat-result.ts Extends integration result types.
apps/brunch-agent/test/local-dev-origins.test.ts Verifies launcher dependencies and API plugins.
apps/brunch-agent/test/architecture/boundaries.integration.ts Registers the new package boundary.
apps/brunch-agent/petrinaut-local.vite.config.ts Preserves website plugins while merging local config.
.changeset/flue-voice-safety.md Records the Petrinaut patch release.
Review details
  • Files reviewed: 79/79 changed files
  • Comments generated: 0
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@codspeed-hq

codspeed-hq Bot commented Sep 7, 2026

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ 6 benchmarks measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. These results are not comparable, so they count as unchanged.

Preventing compiler optimizations

✅ 98 untouched benchmarks

Performance Changes

Benchmark BASE HEAD Efficiency
⚠️ as_constant < 1 ns < 1 ns N/A
⚠️ constant_equal < 1 ns < 1 ns N/A
⚠️ constant_not_equal < 1 ns < 1 ns N/A
⚠️ access < 1 ns < 1 ns N/A
⚠️ runtime_equal < 1 ns < 1 ns N/A
⚠️ runtime_not_equal < 1 ns < 1 ns N/A

Comparing ln/fe-1580-reconcile-voice-resumable-workpiece (743c3c8) with main (6e561a2)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (8b45b5a) during the generation of this report, so 6e561a2 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report.

@lunelson
lunelson deployed to pull-request September 7, 2026 16:11 — with GitHub Actions Active
Base automatically changed from ln/fe-1575-resumable-workpiece-petrinaut to main September 8, 2026 14:53
@lunelson
lunelson force-pushed the ln/fe-1580-reconcile-voice-resumable-workpiece branch from bfd99d3 to b1bb132 Compare September 8, 2026 14:53

@kostandinang kostandinang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved for the narrowed reconciliation scope

This does not authorize retiring #9531 or imply #9538/#9550 are included, for now, until further checks.

The later silent-response fix 9415e1b, interruption/medium-VAD wor work remain to be reconciled separately

@lunelson
lunelson added this pull request to the merge queue Sep 8, 2026
Merged via the queue into main with commit fb96f21 Sep 8, 2026
205 checks passed
@lunelson
lunelson deleted the ln/fe-1580-reconcile-voice-resumable-workpiece branch September 8, 2026 16:30
@hash-release hash-release Bot mentioned this pull request Sep 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/apps > hash.design Affects the `hash.design` design site (app) area/apps area/infra Relates to version control, CI, CD or IaC (area) area/libs Relates to first-party libraries/crates/packages (area) area/tests New or updated tests type/eng > frontend Owned by the @frontend team

Development

Successfully merging this pull request may close these issues.

3 participants