Skip to content

FE-1630: Optimize and measure the Brunch Voice relay - #9585

Open
kostandinang wants to merge 55 commits into
mainfrom
kostandin/fe-1630-improved-voice-relay
Open

FE-1630: Optimize and measure the Brunch Voice relay#9585
kostandinang wants to merge 55 commits into
mainfrom
kostandin/fe-1630-improved-voice-relay

Conversation

@kostandinang

@kostandinang kostandinang commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

🌟 What is the purpose of this PR?

Determine whether Voice-mode prompting and bounded Realtime delivery make the Brunch relay sufficiently natural before reconsidering split ownership.

Implemented experiment; negative naturalness result. Long reports now stay complete on screen and are read only on request. The recorded short clarification is still 151 words and takes 7.4 seconds after transcription to reach a delivery notice, not its answer. Recommendation: the relay still fails for short local follow-ups requiring Brunch round trips; reconsider #9571 using this evidence. This does not approve or implement that architecture.

Kostandin explicitly approved a maintained local Flue 2.0.3 context patch, not an upstream-supported API. The patch exception and removal path are documented. This draft includes real local before/after evidence and an audible demo. After #9564 merged, preview inspection found a Voice config HTTP 500 also present on the parent preview; full preview verification and human acceptance remain outstanding.

🔗 Related links

🚫 Blocked by

  • FE-1580: Reconcile Voice turn behavior on the shared Brunch conversation #9564 merged before any preview inspection.
  • Child ancestry reconciled with current main at ef0f444987; the repository mission conflict retained main’s live authority, and all 258 targeted tests passed again. No parent/sibling or remote history was rewritten.
  • Full preview verification, including the patched Brunch backend; local evidence does not establish that deployment.
  • Owner review of the negative result and demonstration. The experiment’s historical branch contract is not accepted.

The earlier provider credential/routing and per-turn-context prerequisites are resolved: the corrected /agents/chat baseline preceded prompt changes, and the local dependency-patch exception was authorized and committed separately.

🔍 What does this change?

  • Adds optional per-delivery JSON context to pinned Flue runtime/SDK 2.0.3 through Yarn patches, using existing submission/private records and recovery. No new store or provenance claim.
  • Carries responseMode: "voice" through existing Voice user and browser-tool result admissions; injects fixed Brunch system instructions for that delivery only. Typed prompts remain unchanged.
  • Uses a verbatim Realtime renderer with no domain tools or autonomous VAD responses. A long response receives only an application-selected fixed offer, bounded to 256 output tokens, with distinct bridging diagnostics. Requested reading submits unchanged canonical text.
🏗️ Agent notes

The historical experiment contract records the authorized scope without replacing current main’s MISSION.md; the accepted Mission 6b text is preserved verbatim in its archive. No future draft was consumed.

Imperative: Determine whether Voice-mode Brunch prompting and bounded Realtime delivery make the existing relay sufficiently natural without transferring domain authority. Deliver implementation, comparable before/after findings, and a short demonstration video. Concise clarification remains a failed experimental outcome, not a completed acceptance leaf.

Throughline: Existing panel Voice input → shared AI SDK transport → one Flue user admission with unchanged text and optional context.responseMode: "voice" → Brunch fixed effective-system-context instruction → canonical visible response → bounded Realtime speech. Typed messages omit context. Automatic static-browser continuations inherit the originating live preference; dynamic answers carry their own source. No global mode, additional signal admission, or invented user authorship.

Proof: Real local baseline: 192-word clarification with an additional 12-word Realtime preamble; automatic 1,178-word report reading. Final same-input synthetic-speech run: 151-word clarification, 952-word complete visible report, exactly two fixed offers, full canonical text submitted only after Read full response, and Your turn returning to Listening. Actual typed admission omitted context; Stop persisted an aborted Flue settlement; two reloads caused no POST/autoplay or duplicate user turn. Deterministic tests establish context isolation/recovery/idempotency and bounded-delivery lifecycle. These observations are neither a statistical campaign nor human naturalness acceptance; interrupted prefix fidelity is not a full acoustic report test.

Constraints: Brunch retains domain meaning, questions, conclusions, workpiece state, and tools. Realtime has no domain tools, autonomous semantic-VAD response, or authority. Bridging never interprets evidence, confirms a change, asks a domain follow-up, alters qualifications, or invokes tools. Preserve canonical text, shared routing, admission/correlation, interruption versus Stop, exact reading, tools, and reopen. No delegation, clarify_by_voice, client handback, sidecar storage, persistence-only extension, or workpiece/provenance redesign. No unrelated Linear writes, paid evaluation campaign, preview before parent merge, or #9571 edits.

Fog-line: The local patch passes real admission/restart and pinned joined-input recovery tests; it is not upstream support or a killed-process campaign. Fixed instructions did not reliably shorten Brunch output. The final fixed offers matched exactly, but provider compliance is not guaranteed. An initial tool-continuation stall remains unreproduced in the successful fresh run, not declared fixed. Human adequacy, first-audible latency, and architecture choice remain unresolved. The 120-word/1,200-character/fenced-code threshold is experimental. Already-completed short steps may speak before a later long continuation appears.

Stop or reorient: Stop if context leaks, retries duplicate admission, recovery loses preference, typed behavior inherits Voice style, or the patch needs another store/authority. Reorient on observed delivery failures rather than expanding ownership. This packet recommends reconsidering local-follow-up ownership based on the failed experience; it does not implement the future design. If the parent changes, restack/re-pin and rerun affected proof before review. Mission closure remains owner-held.

Dependency contract, owner approval, product implementation, and evidence are separated in local commits. CORS/donor/ownership worktrees were untouched. One FE/brunch-agent issue was created and assigned to Kostandin; no further Linear write was made. Initial environment failures and baseline chronology remain in evidence rather than being erased.

Pre-Merge Checklist 🚀

🚢 Has this modified a publishable library?

This PR:

  • does not modify any publishable blocks or libraries, or modifications do not need publishing

📜 Does this require a change to the docs?

The changes in this PR:

  • are internal and do not require a docs change

Internal experiment/patch documentation and evidence are included. No public Petrinaut package behavior or guide surface was changed.

🕸️ Does this require a change to the Turbo Graph?

The changes in this PR:

  • do not affect the execution graph

⚠️ Known issues

  • Preview blocked: after the implementation's Vercel build succeeded, /api/voice/config returned HTTP 500 (FUNCTION_INVOCATION_FAILED) on both this preview and FE-1580: Reconcile Voice turn behavior on the shared Brunch conversation #9564's preview. The editor/AI shell renders, but Voice controls are absent. The API entrypoint/config handler are unchanged. Hash Vercel function-log access is unavailable to the current account; no environment or deployment was manually changed. The patched backend's remote deployment also remains unverified.
  • Benchmark baseline blocker: the Yarn patch globally invalidates Turbo, selecting @rust/hash-graph-benches. Its base checkout at current main exits before benchmarks because @rust/hash-graph-atlas runs without its required cli feature. This must be repaired on main before this PR’s benchmark can pass.
  • Short Voice clarification still fails: 151 words, repeated explanations, and literal marker markup. Only the delivery notice is spoken because the answer exceeds the budget. The relay preserves rather than silently repairs Brunch content.
  • Long offer latency was 34.7 seconds after transcription. A fixed notice does not remove Brunch round trips. Recorded event timestamps are not first-audible latency metrics.
  • An initial after run stalled after a browser-tool call without a result admission; the final fresh run completed that continuation. The cause remains unresolved; no parent tool-order fix was broadened.
  • A 128-token offer once ended incomplete / max_output_tokens; audio tokens exhausted the budget. The corrected 256-token cap passed final offers at 117/143 tokens, but remains a finite provider-dependent bound.
  • Domain claims in the canonical report were not validated. Full uninterrupted report audio and human conversational acceptance were not tested. The transient Stop notice is not shown after reload; the persisted Flue settlement is aborted. Parent hydration/provenance limitations remain inherited.
  • Initial transitive builds lacked generated graph dependencies. The dependency-patch commit's unrelated Rust hook exhausted disk and was excluded for that commit; other hooks ran. No full-monorepo green claim.

🐾 Next steps

Review the demonstration and failed clarification result. Complete the deployment-specific preview verification now that #9564 has merged. Reconsider #9571 with these findings, without treating this experiment as architecture approval.

🛡 What tests cover this?

  • Brunch targeted runtime, built chat, and architecture: 33 tests / 4 files.
  • Transport: 49 tests / 4 files.
  • Website Voice/policy: 176 tests / 10 files.
  • Builds and TypeScript checks passed for the three affected workspaces. Lint: zero errors, two existing transport warnings and one existing website warning.
  • Changed-file Oxfmt and git diff --check outside the version-pinned patch fixtures passed; those fixtures preserve upstream tab-indented patch context. Browser/real-provider evidence covers long-report withholding, opt-in reading, interruption, typed admission, durable Stop, and reopen.
  • Post-merge preview deployment built successfully. Browser inspection confirmed its shell and the inherited Voice configuration failure; this is not a passing end-to-end preview test.
  • Passing deterministic checks do not mean the short-answer naturalness criterion passed. Commands, precise results, failures, and limitations are in the evidence file.

❓ How to test this?

  1. Use the ignored root .env.local with ANTHROPIC_API_KEY, OPENAI_VOICE_API_KEY, PETRINAUT_OPENAI_VOICE_ENABLED=true, and VITE_BRUNCH_CHAT_ENDPOINT=/agents/chat. Never put a provider key in a VITE_ variable.
  2. Build/start the Brunch app; point BRUNCH_CHAT_ORIGIN at that server when launching yarn workspace @apps/brunch-agent petrinaut:dev with PETRINAUT_WEBSITE_ROOT set to this worktree's website. BRUNCH_DEV_DB_PATH controls local conversation SQLite storage. The recorded final run used server 4324 and panel 4928.
  3. Open /?brunch-fixture=crew-reservation-v1, wait for settled revision zero, open AI, and start Voice after consent. Ask the two exact recorded inputs. Inspect complete canonical text; select the playback menu's Read full response and interrupt with Your turn. Test Stop separately during an admitted response, then reload.
  4. Run the targeted commands in the evidence record. Do not test preview before FE-1580: Reconcile Voice turn behavior on the shared Brunch conversation #9564 merges or interpret this synthetic witness as human acceptance.

📹 Demo

98-second audible demonstration — synthetic Samantha input, real OpenAI and Brunch, inspected audio/video.

demo-2026-09-08.mp4

First offer ≈00:29; long report streams silently ≈00:54–01:22; offer ≈01:22; explicit read ≈01:27; Your turn ≈01:34. The video intentionally retains the failed short-answer outcome.

lunelson and others added 30 commits September 8, 2026 15:55
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…nsport

Prepare one honestly labelled crew-reservation fixture whose starting
workpiece enters canonical Flue history exactly once as a tagged system
dispatch signal with idempotent retry, scope the read/`addArc` client tools
to that fixture, settle a runtime manifest only after conversation,
workpiece, and document state are observable, and surface a visible refusal
when the mounted Brunch route is absent.

The fixture rides Mission 5's landed browser `ChatTransport`: the panel
transport keeps the conversation tracker and accepts fixture-scoped client
tool names as an option, the history hook keeps the live observation and
additionally exposes the observed snapshot with its offset, and AI SDK user
sends carry a deterministic idempotency key.

The real two-tab witness remains blocked on the configured Anthropic
credential; the blocker and handoff are recorded in the FE-1575 evidence
note. No persona testing was run.
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Add dispositions G1 to G22 to the decision log, revise the mini spec to its
third state with the settled-revision protocol, typed basis, verifiable
transition record, document reconciliation, readiness ownership, and probe
decision tables, and retain the follow-up review as design evidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ared basis

Rewrite the Mission 7 draft at cut-level detail as the consolidated
construct-and-explain mission with a two-step authority, probe decision
tables, and a conversion map; recut Mission 9 as repeatable projection
breadth over the Mission 7 seam; re-express Missions 10 and 11 over settled
revisions, declared basis, and transition records with three-gate readiness
for the handoff; and record the provenance relations and tool admission
locks plus a lossless planning-content migration matrix in MISSION.next.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Point the Deferred section at the 2026-09-04 provenance replanning, carry the
fixture-authorship and fenced-block admissions into the close record, and note
in Status that no other section of the live authority changed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Record dispositions H0 to H12 in the decision log, fix revision identity to
the tool call id, make the transition record and elicited evidence relation
precise, replace every oracle gap in the Step A leaves with an exact
prospective test, artifact, or adjudication, classify probe outcomes as
eligible, rework, or terminal stop under the owner's no-narrowing rule, give
the two-step authority a lawful document shape, and add the pre-cut owner
checklist and the ask/sweep keep-remove-archive inventory.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…revision numbering

The index route now owns the brunch-fixture search key beside the shared
example contract, so a selection write no longer drops the prepared
fixture mid-session. The fixture selector and fixture URL fall back to
the ordinary demo while Brunch is unconfigured, the fixture's browser
tool catalog keeps the always-mounted docs reader answerable, workpiece
revisions are numbered by the history rather than by manifest count, and
an unchanged-document mutation result no longer claims the requested
state already existed.
@codspeed-hq

codspeed-hq Bot commented Sep 8, 2026

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ 6 benchmarks measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. These results are not comparable, so they count as unchanged.

Preventing compiler optimizations

✅ 98 untouched benchmarks

Performance Changes

Benchmark BASE HEAD Efficiency
⚠️ as_constant < 1 ns < 1 ns N/A
⚠️ constant_equal < 1 ns < 1 ns N/A
⚠️ constant_not_equal < 1 ns < 1 ns N/A
⚠️ access < 1 ns < 1 ns N/A
⚠️ runtime_equal < 1 ns < 1 ns N/A
⚠️ runtime_not_equal < 1 ns < 1 ns N/A

Comparing kostandin/fe-1630-improved-voice-relay (0902ddd) with main (94dff8e)1

Open in CodSpeed

Footnotes

  1. No successful run was found on main (ef0f444) during the generation of this report, so 94dff8e was used instead as the comparison base. There might be some changes unrelated to this pull request in this report.

@cursor

cursor Bot commented Sep 9, 2026

Copy link
Copy Markdown

PR Summary

Medium Risk
Touches conversation admission/persistence via vendored Flue patches and changes voice UX/policy paths; mitigated by extensive tests but the patch must be re-evaluated on Flue upgrades.

Overview
FE-1630 adds a maintained local Yarn patch to Flue 2.0.3 so user/signal admissions can carry optional JSON context (validated, stored privately, restored via useDelivery(), excluded from public history and model input). Voice traffic sends { responseMode: "voice" } through the AI SDK transport; ChatAgent applies fixed per-delivery voice presentation instructions only when that flag is set.

Petrinaut voice relay is tightened: OpenAI Realtime policy moves to brunch-bounded-relay-v4 (verbatim renderer, no domain authority). Long canonical responses are not auto-read; the bridge plays a fixed “read on screen” bridging offer (256-token cap, distinct diagnostics/events) and readFullResponse speaks full canonical text on demand. Short responses still use canonical TTS. Server policy and turn-controller logic treat bridging speech separately from canonical latency metrics.

Tests and evidence cover context isolation, idempotency/SQLite restart, voice prompt scoping, bounded delivery, and before/after synthetic-speech JSON plus experiment docs under improved-voice-relay/.

Reviewed by Cursor Bugbot for commit 0902ddd. Bugbot is set up for automated code reviews on this repo. Configure here.

kostandinang and others added 2 commits September 9, 2026 10:02
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@codecov

codecov Bot commented Sep 9, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 65.93%. Comparing base (04db878) to head (0902ddd).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #9585   +/-   ##
=======================================
  Coverage   65.93%   65.93%           
=======================================
  Files        1886     1886           
  Lines      198341   198341           
  Branches     8236     8236           
=======================================
  Hits       130772   130772           
  Misses      66039    66039           
  Partials     1530     1530           
Flag Coverage Δ
apps.hash-ai-worker-ts 1.99% <ø> (ø)
apps.hash-api 15.57% <ø> (ø)
apps.hash-graph 12.54% <ø> (ø)
blockprotocol.type-system 38.15% <ø> (ø)
local.claude-hooks 0.00% <ø> (ø)
local.harpc-client 51.49% <ø> (ø)
local.hash-backend-utils 3.27% <ø> (ø)
local.hash-graph-sdk 10.02% <ø> (ø)
local.hash-isomorphic-utils 12.22% <ø> (ø)
rust.antsi 2.36% <ø> (ø)
rust.error-stack 90.81% <ø> (ø)
rust.harpc-codec 84.70% <ø> (ø)
rust.harpc-net 96.21% <ø> (ø)
rust.harpc-tower 67.03% <ø> (ø)
rust.harpc-types 0.00% <ø> (ø)
rust.harpc-wire-protocol 92.23% <ø> (ø)
rust.hash-codec 72.76% <ø> (ø)
rust.hash-config 81.14% <ø> (ø)
rust.hash-graph-api 19.71% <ø> (ø)
rust.hash-graph-atlas 80.36% <ø> (ø)
rust.hash-graph-authentication 96.02% <ø> (ø)
rust.hash-graph-authorization 63.14% <ø> (ø)
rust.hash-graph-embeddings 91.88% <ø> (ø)
rust.hash-graph-postgres-store 32.15% <ø> (ø)
rust.hash-graph-store 48.41% <ø> (ø)
rust.hash-graph-temporal-versioning 50.18% <ø> (ø)
rust.hash-graph-types 0.00% <ø> (ø)
rust.hash-graph-validation 84.71% <ø> (ø)
rust.hash-middleware 90.92% <ø> (ø)
rust.hashql-ast 89.63% <ø> (ø)
rust.hashql-compiletest 28.39% <ø> (ø)
rust.hashql-core 78.92% <ø> (ø)
rust.hashql-diagnostics 72.51% <ø> (ø)
rust.hashql-eval 79.82% <ø> (ø)
rust.hashql-hir 89.09% <ø> (ø)
rust.hashql-mir 87.92% <ø> (ø)
rust.hashql-syntax-jexpr 94.04% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit fd47c96. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/apps area/deps Relates to third-party dependencies (area) area/infra Relates to version control, CI, CD or IaC (area) area/libs Relates to first-party libraries/crates/packages (area) area/tests New or updated tests type/eng > frontend Owned by the @frontend team

Development

Successfully merging this pull request may close these issues.

2 participants