Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 8 additions & 3 deletions docs/spec/03-runtime/02-agent-runtime.md
Original file line number Diff line number Diff line change
Expand Up @@ -959,9 +959,14 @@ they can end the parent turn even if a delegate ignores its abort. Ending the
parent loop does not otherwise abort delegates.

Fatal provider/stream errors (including exhausted HTTP 429) and parent aborts
retain their existing `failed` and `aborted` outcomes. A terminal parent error
also aborts leftover delegates, skips the resume prompt, and returns the
session to idle so Continue is not `AGENT_BUSY` (D352).
retain their existing `failed` and `aborted` outcomes. If an assistant response
ends at the provider's output-token limit (`stopReason: "length"` or
`"max_tokens"`) after emitting report text, the delegate instead settles as
`failed` with `SUBAGENT_OUTPUT_TRUNCATED` and `outputTruncated: true`; its
bounded partial report remains under the failure explanation for diagnosis. A
later delegate turn that ends normally clears the marker and can complete. A
terminal parent error also aborts leftover delegates, skips the resume prompt,
and returns the session to idle so Continue is not `AGENT_BUSY` (D352).

**Resumable delegations (ADR 0279).** `Task` accepts an optional `resume`
parameter carrying the `delegationId` of a settled delegation in the same
Expand Down
1 change: 1 addition & 0 deletions docs/spec/03-runtime/08-error-codes.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,6 +113,7 @@ does not turn temporary thread pressure into a host process exit.
| `SUBAGENT_IDLE_TIMEOUT` | no | withdrawn (D328): idle watchdogs are not armed; the code remains for stored results |
| `SUBAGENT_DURATION_TIMEOUT` | no | withdrawn (D328): duration watchdogs are not armed; the code remains for stored results |
| `SUBAGENT_CONTEXT_OVERFLOW` | no | a delegate's own model context exceeded its safe budget and neither automatic turn-boundary compaction nor the degraded retry that keeps only the task brief and the most recent messages brought it back below the limit; the failure names the actionable recovery instead of the provider's overflow text |
| `SUBAGENT_OUTPUT_TRUNCATED` | no | a delegate's report ended at the model output-token limit; the partial report is preserved for diagnosis, but the run is failed rather than presented as a completed delegation |
### 3.3 Workspace / tools / permissions

| code | retriable | meaning |
Expand Down
23 changes: 23 additions & 0 deletions docs/spec/06-delivery/04-e2e-test-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -10876,6 +10876,29 @@ This test plan spec is accepted when:
- **Milestone**: M6+
- **Status**: Draft. Required suites: `test:e2e`, `test:e2e:subagents`.

#### E2E-SUBAGENT-output-token-limit-is-a-visible-failure

- **Preconditions**: An Agent session uses a deterministic local provider whose
response ends with `stopReason: "length"` or `"max_tokens"` after emitting
non-empty assistant text. The delegate has a valid report and no pending
tool call.
- **Steps**: 1) Delegate the task and let the provider end at its output-token
limit. 2) Read the Task result, lifecycle details, and delegation card. 3)
Resume or retry the same work with a provider response that ends normally.
- **Expected**: The first run settles as `failed`, not `completed`, with
`SUBAGENT_OUTPUT_TRUNCATED` and `outputTruncated: true`. The parent receives
an explicit explanation and the bounded partial report under `Its last
output was:`. A later turn that ends with `stop` clears the marker and
settles as `completed`; the partial run never masquerades as a finished
report.
- **Specs linked**: `03-runtime/02-agent-runtime.md` §5f,
`03-runtime/08-error-codes.md` §3.2
- **Acceptance**: C (conversation), H (diagnostics), Quality
- **Milestone**: M6+
- **Status**: Unit covered by `packages/agent-runtime/src/subagent.test.ts`;
the full desktop journey remains required: `test:e2e`,
`test:e2e:subagents`.

#### E2E-SUBAGENT-context-overflow-reports-actionable-failure

- **Preconditions**: The same injected small-window fake provider, sized so
Expand Down
5 changes: 4 additions & 1 deletion docs/zh-CN/spec/03-runtime/02-agent-runtime.md
Original file line number Diff line number Diff line change
Expand Up @@ -695,7 +695,10 @@ Stop / 运行时销毁。主 Agent 用 `TaskStop` 判断要不要取消;运行
回合打开,等委托完成后再把报告塞回父级。父级收工不会中止它们。

致命的 provider/stream 错误(包括耗尽的 HTTP 429)、父级中止,仍分别保留它们既有的
`failed` 和 `aborted` 结果。
`failed` 和 `aborted` 结果。如果助手响应在提供程序输出 token 上限处结束(`stopReason: "length"`
或 `"max_tokens"`),且已经产生报告文本,该委派会以 `failed`、
`SUBAGENT_OUTPUT_TRUNCATED` 和 `outputTruncated: true` 结算;有界的部分报告会保留在失败说明
下,供诊断截断原因。后续以正常原因结束的委派回合会清除该标记并可以成功完成。
父级终态错误还会中止残留委托、跳过续跑提示,并把会话恢复为空闲,这样
“继续”不会变成 `AGENT_BUSY`(D352)。

Expand Down
1 change: 1 addition & 0 deletions docs/zh-CN/spec/03-runtime/08-error-codes.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,6 +114,7 @@ stdio 与 Tokio 的动态阻塞池隔离,因此后一种情况
| `SUBAGENT_IDLE_TIMEOUT` | 不 | 已撤回(D328):空闲看门狗不再武装;代码仅为已存储结果保留 |
| `SUBAGENT_DURATION_TIMEOUT` | 不 | 已撤回(D328):时长看门狗不再武装;代码仅为已存储结果保留 |
| `SUBAGENT_CONTEXT_OVERFLOW` | 不 | 委派自身的模型上下文超出其安全预算,自动的回合边界压缩与仅保留任务简报和最近消息的降级重试都没能把它带回限制以内;该失败给出可执行的恢复方式,而不是提供商的溢出文本 |
| `SUBAGENT_OUTPUT_TRUNCATED` | 不 | 委派报告在模型输出 token 上限处结束;保留部分报告用于诊断,但该运行报告为失败,不会被呈现为已完成的委派 |

### 3. 3 工作空间/工具/权限

Expand Down
16 changes: 16 additions & 0 deletions docs/zh-CN/spec/06-delivery/04-e2e-test-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -4539,6 +4539,22 @@ eleven-tool-round desktop paths are verified by
- **里程碑**:M6+
- **状态**:草稿。必需套件:`test:e2e`、`test:e2e:subagents`。

#### E2E-SUBAGENT-output-token-limit-is-a-visible-failure:输出 token 上限必须可见地失败

- **先决条件**:Agent 会话使用确定性的本地提供程序;提供程序在产生非空的助手文本后,以
`stopReason: "length"` 或 `"max_tokens"` 结束。委派有有效报告且没有待处理的工具调用。
- **步骤**:1)委派任务并让提供程序在输出 token 上限处结束。2)读取 Task 结果、生命周期
details 和委派卡片。3)用一次正常结束的提供程序响应恢复或重试同一工作。
- **预期**:第一次运行以 `failed` 结算,而不是 `completed`,并带有
`SUBAGENT_OUTPUT_TRUNCATED` 和 `outputTruncated: true`。父级收到明确解释,且部分报告保留在
`Its last output was:` 下面。后续以 `stop` 正常结束的回合会清除标记并以 `completed` 结算;
部分报告的运行绝不会伪装成已完成报告。
- **链接规格**:`03-runtime/02-agent-runtime.md` §5f、`03-runtime/08-error-codes.md` §3.2
- **验收**:C(对话)、H(诊断)、品质
- **里程碑**:M6+
- **状态**:由 `packages/agent-runtime/src/subagent.test.ts` 单元覆盖;完整桌面旅程仍需运行:
`test:e2e`、`test:e2e:subagents`。

#### E2E-SUBAGENT-context-overflow-reports-actionable-failure

- **先决条件**:同一个注入的小窗口伪提供商,窗口小到连降级后的上下文也放不下。
Expand Down
56 changes: 56 additions & 0 deletions packages/agent-runtime/src/subagent.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -367,6 +367,62 @@ describe("SubagentRun reporting", () => {
expect(result.error?.code).toBe("SUBAGENT_NO_REPORT");
});

it("reports truncated output as a failure rather than a clean completion", async () => {
const { run } = createRun();
run.handleEvent({
type: "message_end",
message: assistantMessage({
content: [{ type: "text", text: "Analysis report cut off mid-sentence..." }],
stopReason: "length",
}),
});
run.agent = {
prompt: vi.fn().mockResolvedValue(undefined),
waitForIdle: vi.fn().mockResolvedValue(undefined),
abort: vi.fn(),
};

const result = await (run as unknown as SubagentRun).run();

expect(result.status).toBe("failed");
expect(result.error?.code).toBe("SUBAGENT_OUTPUT_TRUNCATED");
expect(result.outputTruncated).toBe(true);
expect(result.report).toContain("The explorer subagent failed");
expect(result.report).toContain(
"The subagent response exceeded the model's output token limit and was truncated.",
);
expect(result.report).toContain("Analysis report cut off mid-sentence...");
});

it("clears truncated output state if a subsequent turn finishes cleanly", async () => {
const { run } = createRun();
run.handleEvent({
type: "message_end",
message: assistantMessage({
content: [{ type: "text", text: "Initial attempt" }],
stopReason: "length",
}),
});
run.handleEvent({
type: "message_end",
message: assistantMessage({
content: [{ type: "text", text: "Complete final report." }],
stopReason: "stop",
}),
});
run.agent = {
prompt: vi.fn().mockResolvedValue(undefined),
waitForIdle: vi.fn().mockResolvedValue(undefined),
abort: vi.fn(),
};

const result = await (run as unknown as SubagentRun).run();

expect(result.status).toBe("completed");
expect(result.outputTruncated).toBeUndefined();
expect(result.report).toBe("Complete final report.");
});

it("returns aborted without prompting when the parent call is already aborted", async () => {
const controller = new AbortController();
controller.abort();
Expand Down
13 changes: 13 additions & 0 deletions packages/agent-runtime/src/subagent.ts
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,8 @@ export type SubagentRunResult = {
contextCompactions?: number;
/** True when the run had to discard working history without a summary. */
contextDegraded?: boolean;
/** True when the delegate response hit the model's output token limit. */
outputTruncated?: boolean;
error?: { code: string; message: string };
};

Expand Down Expand Up @@ -217,6 +219,7 @@ export class SubagentRun {
private readonly opts: SubagentRunOptions;
private currentAssistant?: UiMessage;
private lastReportText = "";
private lastReportTruncated = false;
private turns = 0;
private toolCalls = 0;
private usage?: MessageUsage;
Expand Down Expand Up @@ -375,6 +378,13 @@ export class SubagentRun {
message: "The subagent finished without writing a report.",
});
}
if (this.lastReportTruncated) {
return this.result("failed", this.lastReportText, {
code: "SUBAGENT_OUTPUT_TRUNCATED",
message:
"The subagent response exceeded the model's output token limit and was truncated.",
});
}
return this.result("completed", this.lastReportText);
}

Expand Down Expand Up @@ -648,6 +658,7 @@ export class SubagentRun {
...(this.modelFailures.length ? { modelFailures: [...this.modelFailures] } : {}),
...(this.contextCompactions > 0 ? { contextCompactions: this.contextCompactions } : {}),
...(this.contextDegraded ? { contextDegraded: true } : {}),
...(this.lastReportTruncated ? { outputTruncated: true } : {}),
...(error ? { error } : {}),
};
}
Expand Down Expand Up @@ -826,6 +837,8 @@ export class SubagentRun {
// must not clear the text an earlier turn already produced.
if (content.hasText && content.text.trim() && !failed) {
this.lastReportText = content.text;
this.lastReportTruncated =
stopReason === "length" || stopReason === "max_tokens";
}
if (retryAttempt !== undefined) {
this.currentAssistant = {
Expand Down
1 change: 1 addition & 0 deletions packages/shared/src/errors.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@ describe("result helpers", () => {
const newlyLiveCodes = [
"CONTEXT_COMPACTION_FAILED",
"EMPTY_MODEL_RESPONSE",
"SUBAGENT_OUTPUT_TRUNCATED",
"EDIT_TAG_REQUIRED",
"EDIT_TAG_MISMATCH",
"EDIT_TAG_UNKNOWN",
Expand Down
1 change: 1 addition & 0 deletions packages/shared/src/errors.ts
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,7 @@ export const ErrorCodes = {
* delegate's model, or how much it reads at once has to change.
*/
SUBAGENT_CONTEXT_OVERFLOW: "SUBAGENT_CONTEXT_OVERFLOW",
SUBAGENT_OUTPUT_TRUNCATED: "SUBAGENT_OUTPUT_TRUNCATED",
WORKSPACE_REQUIRED: "WORKSPACE_REQUIRED",
PATH_OUTSIDE_WORKSPACE: "PATH_OUTSIDE_WORKSPACE",
TOOL_NOT_FOUND: "TOOL_NOT_FOUND",
Expand Down
Loading