Conversation
…onfinement in prompts Instruct agent and subagent runtimes that file contents and tool outputs are untrusted data, and prohibit invoking tools outside the active list. Addresses vastsa#1174
|
I am not merging this PR as a fix for #1174. That issue is already closed by #1184, which changed the runtime boundary itself: historical mode-incompatible tool calls remain declared and are rejected before host execution, with regression coverage in |
|
Thanks for the thorough review and clarification regarding #1184. That makes complete sense — #1184 indeed addressed the programmatic rejection of mode-incompatible tool calls at the gateway level. I've retargeted this PR purely as a defense-in-depth security hardening measure:
If you find this defense-in-depth layering appropriate for the prompt runtime, it's ready for review. Otherwise, feel free to close if you prefer keeping prompts minimal. Thanks again! |
`main` is currently red, and it is not a flake:
CI vastsa#1197 166ada6 failure Typecheck Desktop
CI vastsa#1195 967f3c2 failure Typecheck Desktop
CI vastsa#1192 db16f9c failure Typecheck Desktop
CI vastsa#1189 4fedb8e success <- last green main run I can see
All three failures are the same line:
electron/main/user-login-path.ts(28,74): error TS2345: Argument of type
'string | undefined' is not assignable to parameter of type 'PathLike'.
**Root cause.** `Boolean(candidate) && existsSync(candidate)` is inside a type
predicate:
const shell = [process.env.SHELL, "/bin/zsh", "/bin/bash", "/bin/sh"].find(
(candidate): candidate is string => Boolean(candidate) && existsSync(candidate),
);
A type predicate constrains what the *caller* sees; it does not narrow the
parameter inside its own body, and `Boolean(candidate)` is not a narrowing call.
So the body still has `string | undefined` and `existsSync` rejects it.
**Fix.** Narrow with a real check, which keeps the predicate at the call site and
satisfies the body:
(candidate): candidate is string =>
typeof candidate === "string" && existsSync(candidate),
The file arrived with `574c78a9` ("fix(mcp): spawn stdio servers with the
login-shell PATH"). That commit is 1 ahead / 8 behind `4fedb8e4` — it landed on
`main` after the last green run, which is where main turns red.
**Verification** (Windows 11; `pnpm install --filter @pi-desktop/desktop...
--ignore-scripts`, `pnpm --filter "@pi-desktop/desktop^..." build`):
- Pristine file, the CI command `pnpm --filter @pi-desktop/desktop typecheck`
→ exit 2 with `user-login-path.ts(28,74): error TS2345` — the same file,
line and column as CI. With the fix → **exit 0**.
- `pnpm lint` → exit 0.
- `node scripts/check-architecture.mjs` → Architecture check passed.
- `node --test test/user-login-path.test.mjs` → 4/4 pass.
- Full `apps/desktop` suite: 2452 tests; 43 fail with the pristine file and 42
with the fix. Both failure sets are the same macOS-only group (codesign, DMG,
notarization, ssh askpass, process groups) plus a couple of load-sensitive
timeout-budget tests; I could not make those two pass on either revision, and
nothing in this change can reach them. Recording it rather than claiming the
suite is green.
|
Closing to keep upstream scope clean, as this is proactive hardening rather than a root fix for an open bug. We can revisit with a dedicated issue if prompt-level boundaries are desired later. |
Summary
Adds explicit untrusted-data boundary instructions to the main agent and subagent operational system prompts as defense-in-depth security hardening against indirect prompt injection (IPI).
Motivation & Security Acceptance
While repository guidelines (
AGENTS.md§11) specify that repository content, issue text, tool outputs, and user files are passive data rather than instructions, this rule was missing from operational runtime prompts (defaultSystemPromptPartsandcomposeSubagentSystemPrompt). Models reading external files or untrusted tool outputs could be vulnerable to embedded prompt overrides trying to redirect execution or escalate capabilities.This change introduces prompt-level defense-in-depth to explicitly remind the model to treat external inputs as passive data.
Key Changes
packages/agent-runtime/src/runtime.ts:defaultSystemPromptParts:packages/agent-runtime/src/subagent.ts:composeSubagentSystemPrompt.packages/agent-runtime/src/runtime.test.ts&subagent.test.ts:Verification
runtime.test.ts: 291/291 passed.subagent.test.ts: passed.