Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
101 commits
Select commit Hold shift + click to select a range
91e28ce
Cut Mission 4: redesign runbook and workpiece
lunelson Aug 31, 2026
7347f6f
Advance upcoming missions after Mission 4 cut
lunelson Aug 31, 2026
3a87a60
Keep closed Mission 3 outside future numbering
lunelson Aug 31, 2026
41a1214
Bind Mission 4 to the completed baseline
lunelson Aug 31, 2026
237a311
Colocate the chat agent with its Flue resources
lunelson Aug 31, 2026
93d81db
Group Brunch shell code by authority
lunelson Aug 31, 2026
d090c2b
Separate runbook instruments from product runtime
lunelson Aug 31, 2026
d2888af
Keep Brunch topology verification current
lunelson Aug 31, 2026
767c326
Inject the in-process Flue transport
lunelson Aug 31, 2026
2d6d0dd
Split Brunch core from SDCPN agent resources
lunelson Aug 31, 2026
3ada120
Inline short Flue prompt fragments
lunelson Aug 31, 2026
42a78c7
Format Flue instructions as template blocks
lunelson Aug 31, 2026
eb58763
Keep Brunch core context independent
lunelson Aug 31, 2026
c8ab012
Audit generic elicitor prompt material
lunelson Aug 31, 2026
70acc0d
Load the Brunch system prompt from Markdown
lunelson Aug 31, 2026
96479d7
Add a core prompt shaping workbench
lunelson Aug 31, 2026
4f5a1e6
Group Brunch core source by authority
lunelson Aug 31, 2026
ecf4484
Record the core prompt topology decision
lunelson Aug 31, 2026
aff92de
WIP
lunelson Aug 31, 2026
22613c0
Keep chat verification aligned with core identity
lunelson Sep 1, 2026
380262d
Test chat behavior instead of prompt wording
lunelson Sep 1, 2026
59d22b0
WIP
lunelson Aug 31, 2026
3301937
pre-consultation
lunelson Sep 1, 2026
61a256f
new full pattern drafts
lunelson Sep 1, 2026
75b7abe
two good draft prompting sets; domain-typology and target-formalism a…
lunelson Sep 1, 2026
4719848
Refine the five-register elicitation draft
lunelson Sep 1, 2026
0862554
Calibrate evidence claims in the runbook draft
lunelson Sep 1, 2026
3d89fee
Define the five-register draft evaluation protocol
lunelson Sep 1, 2026
65085cb
Generalize the Ampcode prompt architecture draft
lunelson Sep 1, 2026
d08d5ad
Add the domain-primary readiness candidate
lunelson Sep 1, 2026
ee8d01a
Record the five-register mechanical audit
lunelson Sep 1, 2026
302e3f5
five-register post-review fixing
lunelson Sep 1, 2026
a255878
Archive Mission 4 at branch transition
lunelson Sep 1, 2026
c10c359
Record Flue skill composition evidence
lunelson Sep 2, 2026
f179ee7
side quest for mission re-lacing
lunelson Sep 2, 2026
ff33ff4
mission re-lacing sidequest after 3 passes
lunelson Sep 2, 2026
d14b461
Restore Brunch planning decisions
lunelson Sep 2, 2026
a72cd6f
Pin Brunch planning split proof
lunelson Sep 2, 2026
819268b
Freeze the Brunch Mission 4 candidate instrument
lunelson Sep 2, 2026
702027b
Narrow Mission 4 scoring to the flat-prompt control
lunelson Sep 2, 2026
79529cb
Close Brunch Mission 4 with accepted evidence
lunelson Sep 2, 2026
2d09a34
ignore playwright cli outputs
lunelson Sep 2, 2026
da9f926
Reopen and repair Brunch Mission 4 evidence
lunelson Sep 2, 2026
ed61aeb
Freeze the repaired Mission 4 campaign protocol
lunelson Sep 2, 2026
df5384b
Authorize the repaired Mission 4 scoring campaign
lunelson Sep 2, 2026
04781b6
manual docs move
lunelson Sep 2, 2026
31def20
Reject the invalid Mission 4 v4 campaign
lunelson Sep 2, 2026
f5b7302
Freeze the Mission 4 v5 routing repair
lunelson Sep 2, 2026
7210b66
Lay the agreed Brunch core/plugin topology and retire the YAML plugin…
lunelson Sep 2, 2026
975cca9
draft material deletions;
lunelson Sep 2, 2026
0877e27
doc salvage addition
lunelson Sep 2, 2026
febb173
Adopt decision-integrity rules and preserve the Mission 4 design record
lunelson Sep 2, 2026
018bb2a
Recut Mission 4 as the accepted core/plugin architecture proof
lunelson Sep 2, 2026
b1f4697
Add "brunch-turn" local pi extension for testing
lunelson Sep 2, 2026
829c477
Improve Brunch persona harness controls
lunelson Sep 2, 2026
a6397e9
Deduplicate Pi TUI lockfile dependency
lunelson Sep 2, 2026
a7d6015
Add client-tool hosting to Brunch persona harness
lunelson Sep 2, 2026
ccaa3d5
Add SDCPN common-language reference
lunelson Sep 2, 2026
5476443
Model persona interaction posture explicitly
lunelson Sep 2, 2026
3576da5
Add prospective persona evaluation cases
lunelson Sep 2, 2026
dc37b36
Move the Brunch persona harness into the application
lunelson Sep 2, 2026
7cd8788
Drop the duplicate SDCPN common-language salvage copy
lunelson Sep 2, 2026
ba31a6c
Propose the Mission 4 activation and restraint ruler
lunelson Sep 2, 2026
3d7cc94
mission 6 addition re petrinaut tool
lunelson Sep 2, 2026
1f7f3f4
Accept the Mission 4 proof-of-life ruler
lunelson Sep 3, 2026
4b403d2
Recut Mission 4 around proof of life
lunelson Sep 3, 2026
b3d0947
Inline universal elicitation guidance
lunelson Sep 3, 2026
e6796f6
Plan the Mission 4 proof-of-life campaign
lunelson Sep 3, 2026
1e940d5
Retain canonical persona proof traces
lunelson Sep 3, 2026
3b41199
Bind persona proof artifacts by hash
lunelson Sep 3, 2026
60ebbf1
Record Mission 4 model and cost preflight
lunelson Sep 3, 2026
a589e55
Review the proof trace substrate boundary
lunelson Sep 3, 2026
f90cf72
Prepare the Mission 4 freeze candidate
lunelson Sep 3, 2026
2c55621
Freeze the Mission 4 instrument manifest
lunelson Sep 3, 2026
7a7ee78
Record the Mission 4 freeze authorization
lunelson Sep 3, 2026
d9646f5
Retain the first Vestera proof probe
lunelson Sep 3, 2026
510b6c7
Retain the replacement Vestera proof probe
lunelson Sep 3, 2026
229b8fa
Cut the Mission 4 proof campaign v2
lunelson Sep 3, 2026
d75930b
Harden the Mission 4 v2 freeze candidate
lunelson Sep 3, 2026
9227420
Prepare the Mission 4 v2 freeze manifest
lunelson Sep 3, 2026
d7a93fb
Suspend the Mission 4 v2 currency gate
lunelson Sep 3, 2026
d3ef6d6
Refresh the Mission 4 v2 freeze manifest
lunelson Sep 3, 2026
f4355f5
Record the Mission 4 v2 freeze acceptance
lunelson Sep 3, 2026
a082457
Retain the Mission 4 v2 Vestera probe
lunelson Sep 3, 2026
bfd984e
Retain the Mission 4 v2 Data Centre probe
lunelson Sep 3, 2026
0a6fa30
Retain the Mission 4 v2 S3 review
lunelson Sep 3, 2026
bad62d1
Retain the Mission 4 v2 S4 failure
lunelson Sep 3, 2026
a1b1371
Record the Mission 4 v2 campaign stop
lunelson Sep 3, 2026
1cb5c77
Refresh the Mission 4 v2 run manifests
lunelson Sep 3, 2026
2215c64
Record the Mission 4 v2 proof failure
lunelson Sep 3, 2026
37a0574
Close Mission 4 and hand off Voice integration
lunelson Sep 3, 2026
95aad25
Simplify Mission 4 restack provenance
lunelson Sep 3, 2026
aa034de
Record the Voice merge collision probe
lunelson Sep 3, 2026
daf1d24
Decouple Mission 4 evidence from restack SHAs
lunelson Sep 3, 2026
f313168
Resolve Brunch stack review failures
lunelson Sep 3, 2026
72d51c9
re-write next missions
lunelson Sep 3, 2026
599328f
Link successor missions to Linear issues
lunelson Sep 3, 2026
733f1d0
Retire raw skill-composition run records
lunelson Sep 3, 2026
69c7d35
pm-oriented mission re-drafts
lunelson Sep 3, 2026
2563680
Retain compressed side-quest run evidence
lunelson Sep 3, 2026
51bec22
Align upper Brunch test task dependencies
lunelson Sep 3, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@ bower_components
.lock-wscript

# Tests
**/.playwright-cli/
**/playwright-report/
**/test-results/
tests/**/.auth/
Expand Down
127 changes: 127 additions & 0 deletions apps/brunch-agent/.pi/extensions/brunch-persona-testing.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
/**
* Pi extension entry for the Brunch persona harness.
*
* Load it from `apps/brunch-agent` with
* `--extension .pi/extensions/brunch-persona-testing.ts`; the persona policy
* and operating instructions sit in the same-named folder beside this file.
* This entry owns only Pi registration and flag handling. The tool lives in
* `src/evaluations/persona/brunch-turn.ts` and the client-tool hosts in
* `src/evaluations/persona/client-tool-hosts.ts`, where the application's
* lint, type-check, and unit tests govern them.
*/
import {
type BrunchTurnExtensionApi,
registerBrunchTurn,
requireConversationId,
} from "../../src/evaluations/persona/brunch-turn.ts";
import {
type BrunchClientToolHost,
createMockClientToolHost,
createRealHeadlessClientToolHost,
readMockCalls,
TOOL_HOST_FLAG,
} from "../../src/evaluations/persona/client-tool-hosts.ts";
import { writeProofArtifacts } from "../../src/evaluations/persona/proof-artifacts.ts";

/** The slice of Pi's extension API this entry needs; Pi itself is not a workspace dependency. */
interface BrunchPersonaExtensionApi extends BrunchTurnExtensionApi {
registerFlag(
name: string,
options: {
readonly description?: string;
readonly type: "string";
readonly default?: string;
},
): void;
getFlag(name: string): boolean | string | undefined;
on(
event: "session_start" | "session_shutdown",
handler: () => void | Promise<void>,
): void;
}

const TOOL_MOCKS_FLAG = "brunch-tool-mocks";
const HEADLESS_TITLE_FLAG = "brunch-headless-title";
const EVIDENCE_DIRECTORY_FLAG = "brunch-evidence-dir";

const stringFlag = (
pi: BrunchPersonaExtensionApi,
name: string,
): string | undefined => {
const value = pi.getFlag(name);
return typeof value === "string" && value.trim().length > 0
? value.trim()
: undefined;
};

const createConfiguredClientToolHost = (
pi: BrunchPersonaExtensionApi,
): BrunchClientToolHost | undefined => {
const mode = stringFlag(pi, TOOL_HOST_FLAG) ?? "none";
if (mode === "none") return undefined;

if (mode === "mock") {
const fixturePath = stringFlag(pi, TOOL_MOCKS_FLAG);
if (fixturePath === undefined) {
throw new Error(
`--${TOOL_MOCKS_FLAG} is required when --${TOOL_HOST_FLAG}=mock`,
);
}
return createMockClientToolHost(readMockCalls(fixturePath));
}

if (mode === "real-headless") {
const title =
stringFlag(pi, HEADLESS_TITLE_FLAG) ??
`Brunch persona ${requireConversationId(process.env["PI_SUBAGENT_NAME"])}`;
return createRealHeadlessClientToolHost(title);
}

throw new Error(
`--${TOOL_HOST_FLAG} must be one of none, mock, or real-headless; received ${mode}`,
);
};

// Pi loads an extension through its default export.
export default function brunchPersonaTestingExtension(
pi: BrunchPersonaExtensionApi,
): void {
pi.registerFlag(TOOL_HOST_FLAG, {
type: "string",
default: "none",
description: "Client-tool host: none, mock, or real-headless",
});
pi.registerFlag(TOOL_MOCKS_FLAG, {
type: "string",
description: "Ordered JSON fixture used by the mock client-tool host",
});
pi.registerFlag(HEADLESS_TITLE_FLAG, {
type: "string",
description: "Document title used by the real-headless Petrinaut host",
});
pi.registerFlag(EVIDENCE_DIRECTORY_FLAG, {
type: "string",
description:
"Directory for canonical snapshot, transcript, and trace files",
});

let clientToolHost: BrunchClientToolHost | undefined;
pi.on("session_start", async () => {
await clientToolHost?.dispose?.();
clientToolHost = createConfiguredClientToolHost(pi);
});
pi.on("session_shutdown", async () => {
await clientToolHost?.dispose?.();
clientToolHost = undefined;
});

registerBrunchTurn(pi, {
resolveClientToolHost: () => clientToolHost,
retainSnapshot: async (snapshot) => {
const directory = stringFlag(pi, EVIDENCE_DIRECTORY_FLAG);
if (directory !== undefined) {
await writeProofArtifacts(directory, snapshot);
}
},
});
}
160 changes: 160 additions & 0 deletions apps/brunch-agent/.pi/extensions/brunch-persona-testing/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,160 @@
# Brunch persona testing

This folder holds the persona policy and operating instructions for the local Pi extension at [`../brunch-persona-testing.ts`](../brunch-persona-testing.ts), which drives the production Brunch elicitor as an automated user persona. It records implementation decisions and operating instructions, not execution authority; the Brunch context root's [`MISSION.md`](../../../../../libs/@hashintel/brunch-agent/MISSION.md) remains the live mission when one exists.

## Ownership and layout

The harness is a client of this application's composition: it derives Flue identity through the application's identity authority, resumes Brunch through the application's client-tool signal, and reuses the application's headless Petrinaut client. The application therefore owns it, declares its dependencies, and governs it with its own lint, type-check, and unit tests. It consumes reusable case inputs from the Brunch context root's `evaluations/`.

- [`../brunch-persona-testing.ts`](../brunch-persona-testing.ts) is the Pi entry: registration, flags, and client-tool host selection.
- [`src/evaluations/persona/brunch-turn.ts`](../../../src/evaluations/persona/brunch-turn.ts) owns the `brunch_turn` tool: Flue identity and turn correlation, client-tool resume signals, the evaluation-side tool trace, and rendering.
- [`src/evaluations/persona/client-tool-hosts.ts`](../../../src/evaluations/persona/client-tool-hosts.ts) owns the mock and real-headless client-tool hosts.
- [`SYSTEM.md`](SYSTEM.md) owns only the persona's private policy and epistemic behavior.
- [`src/ui/chat.tsx`](../../../src/ui/chat.tsx) owns the independently attachable read-only browser projection.
- [`test/brunch-turn.test.ts`](../../../test/brunch-turn.test.ts) pins the bridge and tool-host contract.
- [The original spike evidence](../../../../../libs/@hashintel/brunch-agent/docs/evidence/evaluations/live-observable-persona-spike/README.md) records the observed text-only run and proof disposition at the paths used by that run.
- [`MISSION.next.md`](../../../../../libs/@hashintel/brunch-agent/MISSION.next.md#observability-and-simulation-viewing) owns future observability and simulation-viewing work.

Reusable interviewee-visible source truth belongs under
[`evaluations/cases/`](../../../../../libs/@hashintel/brunch-agent/evaluations/cases/), hidden answer keys under
[`evaluations/oracles/`](../../../../../libs/@hashintel/brunch-agent/evaluations/oracles/), and prompts, fixtures, runners, and
procedures under [`evaluations/protocols/`](../../../../../libs/@hashintel/brunch-agent/evaluations/protocols/). Vestera is the
executed exemplar. Industrial-gas VMI, truck-fleet maintenance, semiconductor-fab operations,
data-centre thermal operations, and pharma cold chain now have full greenfield situation packs,
opening messages, and prospective ledgers. They remain unvalidated in 6–10-turn runs and are not
yet bound to a frozen protocol. Bounded approval probes remain skill-composition probes rather
than general persona cases.

## Context and authority boundary

The situation pack, objective, uncertainty, and turn budget are supplied only in the Pi persona's launch prompt. They are not added to the Flue conversation, the Brunch `ChatAgent` instructions, or a Brunch tool payload.

`brunch_turn` accepts exactly one model-authored field:

```ts
brunch_turn({ message: string });
```

It sends that string as one visible Flue user message. The remaining outbound values are guarded conversation identity, incarnation data, andβ€”only after Brunch itself requests a client-deferred toolβ€”the result signal for that call. The raw pack, objective, budget, persona instructions, host configuration, and tool trace have no automatic path into Brunch.

This is prompt-enforced semantic privacy, not formal non-interference. Facts the persona deliberately or accidentally puts in `message` become part of canonical Brunch history. The accepted policy is that `SYSTEM.md` must preserve private instructions and disclose scenario knowledge only through in-character answers; there is no semantic egress filter. The pack is also visible to the persona provider, local Pi process/session, and operator. β€œPrivate” here means isolated from the elicitor context, not secret from that execution environment.

Pi sends only the tool result's `content` back to the persona model. The structured `details` and custom rendering are operator-side metadata, so the tool activity trace does not enter either the persona's next model turn or Brunch's conversation.

```text
private situation pack + objective
β†’ Pi persona chooses one in-character utterance
β†’ brunch_turn sends only that utterance
β†’ production Brunch Flue ChatAgent replies or requests tools
β”œβ”€ server tools execute inside Flue
└─ client-deferred tools execute in the explicitly selected harness host
β†’ one canonical client-tool-result signal resumes Brunch
β†’ Flue stores canonical conversation history
β”œβ”€ Pi renders the actor/process view plus evaluation-side tool activity
β”œβ”€ browser renders a read-only product view
└─ transcript CLI renders the durable audit view
```

The Pi persona is an evaluation-side user actor, not a second Brunch elicitor. Its TUI is an operator harness, not product UI. The browser observer is a local debug projection, not another writer, transcript authority, or inferential observer.

## Transport and identity decisions

- Use the existing mounted production `ChatAgent`; never spawn or emulate another elicitor.
- Use the fixed local principal `local` and `PI_SUBAGENT_NAME` as the conversation id. Derive the Flue instance id and ownership headers through the app's existing identity authority.
- Require a unique, non-empty child name. Never silently generate or switch identity.
- Send the first turn with `uid: null`, then pin every later user or resume send to the returned incarnation `uid`.
- Correlate each response with submission-scoped `read(admission)`. Inspect `history()` only after settlement to find dynamic-tool parts belonging to that submission; never select the latest assistant reply from history.
- Permit one active call at a time and one visible Flue user message per admitted `brunch_turn` call.
- Never resend a user utterance after admission. If settlement, tool hosting, or resume becomes indeterminate, preserve the failure, block later sends from that process, and inspect canonical history.
- Keep the browser observer independently attachable and read-only. Normal local chat remains writable and keeps its generated conversation id.

## Client-tool hosts

`--brunch-tool-host` has three explicit modes:

- `none` is the default. Flue still executes server tools and the Pi result records them. A client-deferred call fails loudly instead of hanging or fabricating a result.
- `mock` consumes an ordered JSON fixture supplied by `--brunch-tool-mocks <path>`, resolved against the working directory. Every tool name and input must match exactly. Missing, extra, or out-of-order calls fail the admitted turn and block later sends.
- `real-headless` executes `readPetrinautDoc` against the checked-out Petrinaut user guide and executes supported construction calls through the existing headless Petrinaut callbacks. `--brunch-headless-title <title>` controls the in-memory document title.

Selecting a host does not mount tools, set Flue initial data, or change production composition. It services only client-deferred calls the real production agent emits. The normal persona route currently mounts `readPetrinautDoc`; construction tools remain conditional on the production agent's validated-construction mode. `real-headless` is real core callback execution against an in-memory document, not browser UI execution, browser rendering, persistence, or proof of product parity.

One suspension may contain multiple calls. The bridge executes them sequentially in canonical order, sends one `client-tool-result` signal carrying their existing call ids, then performs another submission-scoped read. It repeats for at most 20 client-tool rounds and does not return early merely because a suspending response also contained text.

A mock fixture has this shape and should live with the evaluation protocol that owns it:

```json
{
"calls": [
{
"toolName": "readPetrinautDoc",
"input": { "doc": "simulation" },
"output": "Fixture-controlled page text"
}
]
}
```

The final Pi tool details contain every observed server call and every hosted client call with sequence, Flue submission id, tool call id, tool name, executor (`server`, `mock`, or `real-headless`), outcome, input, and output/error. `renderResult` shows a concise `### Tool activity` list beneath `## Brunch`; raw values remain in details and canonical tool activity remains available through the transcript.

When `--brunch-evidence-dir <attempt-directory>` is supplied, every settled `history()` read atomically refreshes `snapshot.json`, `transcript.md`, `trace.json`, `trace.md`, any recovered `workpiece.md` plus `workpiece-source.json`, and `manifest.json` in that directory before `brunch_turn` returns or handles a pending client tool. The snapshot is canonical; transcript, trace, and workpiece recovery are deterministic projections. This retention also occurs before host-none reports an unsupported client-tool suspension.

Create the protocol-owned `run.json` in the attempt directory before launch; it is included in `manifest.json` without being interpreted by the harness. After adding `validity.json` or `adjudication.md`, refresh all sibling hashes with `yarn workspace @apps/brunch-agent proof:manifest -- <attempt-directory>`. Temporary files and `manifest.json` itself are excluded from the manifest.

Pi's tool API requires TypeBox parameter schemas, so `typebox` is declared here for that Pi-facing boundary only. Brunch's own boundaries remain Valibot.

## Operating the harness

1. Start the local app with `yarn workspace @apps/brunch-agent dev`.
2. From `apps/brunch-agent`, launch Pi (directly or through Herdr) with a unique `PI_SUBAGENT_NAME`. Choose the persona model and thinking level with Pi's native `--model <provider/model>` and `--thinking <level>` options.
3. Supply the situation pack inline with the objective and turn budget. The extension treats this launch content as opaque Markdown or plain text and does not parse or validate a pack schema. An `@file` token in a launch task is not expanded into persona context.
For comparable runs, use only the text below the `---` separator in the case's
`opening-message.md` as the visible first turn; keep its header, the situation pack, and the
oracle private. In a 6–10-turn run, bound the objective to the named incident and its immediate
options rather than asking the persona to disclose the entire pack.
4. Wait until the first `brunch_turn` admission is visible in Pi.
5. Attach the browser to `http://127.0.0.1:4321/?mode=observe&principal=local&id=<PI_SUBAGENT_NAME>`.
6. After the run, inspect canonical history with `yarn workspace @apps/brunch-agent transcript -- --principal local --id <PI_SUBAGENT_NAME>`.

The restricted direct launch, run from `apps/brunch-agent`, is:

```sh
PI_SUBAGENT_NAME=<unique-conversation-id> pi \
--model <provider/model> \
--thinking <level> \
--no-extensions \
--extension .pi/extensions/brunch-persona-testing.ts \
--no-builtin-tools \
--tools brunch_turn \
--no-skills \
--no-prompt-templates \
--no-context-files \
--append-system-prompt .pi/extensions/brunch-persona-testing/SYSTEM.md \
--brunch-tool-host real-headless \
--brunch-headless-title "Persona evaluation" \
--brunch-evidence-dir ../../libs/@hashintel/brunch-agent/docs/evidence/evaluations/<campaign>/runs/<attempt-id> \
--approve
```

For deterministic mocks, replace the last host options with:

```sh
--brunch-tool-host mock \
--brunch-tool-mocks ../../libs/@hashintel/brunch-agent/evaluations/protocols/<protocol>/client-tools.json
```

`--no-extensions` plus the one explicit `--extension` prevents dependence on unrelated active Pi extensions. Herdr can forward the same native Pi arguments after `--`; any Herdr companion/state extension is optional orchestration rather than part of the Brunch transport. The persona must never use a parent to obtain domain facts or decide how to answer.

The ordering in steps 4–5 is required by observed behavior. An observer opened before the Flue instance exists remains idle and does not discover later creation. Attaching after first admission catches up existing history and receives later streaming updates. Reloading after creation reconstructs settled messages.

[`evaluations/cases/vestera-scheduling/situation-pack.md`](../../../../../libs/@hashintel/brunch-agent/evaluations/cases/vestera-scheduling/situation-pack.md) is the current exemplar pack. Its Markdown sections are guidance for the persona model, not fields consumed by the extension.

## Rejected alternatives and limits

- A nested `persona β†’ elicitor subagent` topology, because it duplicates elicitor authority and bypasses the real boundary.
- Registering Brunch's tools with Pi, which would expose capabilities to the persona model and move execution to the wrong authority. The host is internal to `brunch_turn` and services only Flue-emitted calls.
- Injecting tool traces or host instructions into Brunch messages. Flue history is canonical; Pi details are an evaluation-side projection.
- Keeping the extension in the Brunch context root, which is not a package: it imported application internals across the lib/app boundary, its dependencies were declared elsewhere or nowhere, and no workspace lint or type-check task reached it.
- `pi-web`, a Herdr webview, PTY scraping, parent-mediated turn relaying, a second server, another model loop, or another transcript store.
- Reply recovery from the latest history entry, automatic user-message retries, pending-admission persistence, or cross-process adoption before a real consumer requires them.

The original evidence establishes a local, live-observable, multi-turn text path through real production code with singular writers, exact submission correlation, browser catch-up/streaming/reload, and transcript parity. The added host modes are covered by local contract tests and extension loading. They do not establish a paid live tool turn, deployed throughline, browser parity, full elicitation quality, persona fidelity across cases, repeatability, crash recovery, pre-creation observer discovery, or remote access.
Loading
Loading