diff --git a/docs/spec/03-runtime/02-agent-runtime.md b/docs/spec/03-runtime/02-agent-runtime.md index 8c499a416..94eacf3de 100644 --- a/docs/spec/03-runtime/02-agent-runtime.md +++ b/docs/spec/03-runtime/02-agent-runtime.md @@ -959,9 +959,14 @@ they can end the parent turn even if a delegate ignores its abort. Ending the parent loop does not otherwise abort delegates. Fatal provider/stream errors (including exhausted HTTP 429) and parent aborts -retain their existing `failed` and `aborted` outcomes. A terminal parent error -also aborts leftover delegates, skips the resume prompt, and returns the -session to idle so Continue is not `AGENT_BUSY` (D352). +retain their existing `failed` and `aborted` outcomes. If an assistant response +ends at the provider's output-token limit (`stopReason: "length"` or +`"max_tokens"`) after emitting report text, the delegate instead settles as +`failed` with `SUBAGENT_OUTPUT_TRUNCATED` and `outputTruncated: true`; its +bounded partial report remains under the failure explanation for diagnosis. A +later delegate turn that ends normally clears the marker and can complete. A +terminal parent error also aborts leftover delegates, skips the resume prompt, +and returns the session to idle so Continue is not `AGENT_BUSY` (D352). **Resumable delegations (ADR 0279).** `Task` accepts an optional `resume` parameter carrying the `delegationId` of a settled delegation in the same diff --git a/docs/spec/03-runtime/08-error-codes.md b/docs/spec/03-runtime/08-error-codes.md index 5a8f80c8e..17190ee47 100644 --- a/docs/spec/03-runtime/08-error-codes.md +++ b/docs/spec/03-runtime/08-error-codes.md @@ -113,6 +113,7 @@ does not turn temporary thread pressure into a host process exit. | `SUBAGENT_IDLE_TIMEOUT` | no | withdrawn (D328): idle watchdogs are not armed; the code remains for stored results | | `SUBAGENT_DURATION_TIMEOUT` | no | withdrawn (D328): duration watchdogs are not armed; the code remains for stored results | | `SUBAGENT_CONTEXT_OVERFLOW` | no | a delegate's own model context exceeded its safe budget and neither automatic turn-boundary compaction nor the degraded retry that keeps only the task brief and the most recent messages brought it back below the limit; the failure names the actionable recovery instead of the provider's overflow text | +| `SUBAGENT_OUTPUT_TRUNCATED` | no | a delegate's report ended at the model output-token limit; the partial report is preserved for diagnosis, but the run is failed rather than presented as a completed delegation | ### 3.3 Workspace / tools / permissions | code | retriable | meaning | diff --git a/docs/spec/06-delivery/04-e2e-test-plan.md b/docs/spec/06-delivery/04-e2e-test-plan.md index 91564c7fe..4363115c4 100644 --- a/docs/spec/06-delivery/04-e2e-test-plan.md +++ b/docs/spec/06-delivery/04-e2e-test-plan.md @@ -10876,6 +10876,29 @@ This test plan spec is accepted when: - **Milestone**: M6+ - **Status**: Draft. Required suites: `test:e2e`, `test:e2e:subagents`. +#### E2E-SUBAGENT-output-token-limit-is-a-visible-failure + +- **Preconditions**: An Agent session uses a deterministic local provider whose + response ends with `stopReason: "length"` or `"max_tokens"` after emitting + non-empty assistant text. The delegate has a valid report and no pending + tool call. +- **Steps**: 1) Delegate the task and let the provider end at its output-token + limit. 2) Read the Task result, lifecycle details, and delegation card. 3) + Resume or retry the same work with a provider response that ends normally. +- **Expected**: The first run settles as `failed`, not `completed`, with + `SUBAGENT_OUTPUT_TRUNCATED` and `outputTruncated: true`. The parent receives + an explicit explanation and the bounded partial report under `Its last + output was:`. A later turn that ends with `stop` clears the marker and + settles as `completed`; the partial run never masquerades as a finished + report. +- **Specs linked**: `03-runtime/02-agent-runtime.md` §5f, + `03-runtime/08-error-codes.md` §3.2 +- **Acceptance**: C (conversation), H (diagnostics), Quality +- **Milestone**: M6+ +- **Status**: Unit covered by `packages/agent-runtime/src/subagent.test.ts`; + the full desktop journey remains required: `test:e2e`, + `test:e2e:subagents`. + #### E2E-SUBAGENT-context-overflow-reports-actionable-failure - **Preconditions**: The same injected small-window fake provider, sized so diff --git a/docs/zh-CN/spec/03-runtime/02-agent-runtime.md b/docs/zh-CN/spec/03-runtime/02-agent-runtime.md index 8ee838d4d..19e78cc1a 100644 --- a/docs/zh-CN/spec/03-runtime/02-agent-runtime.md +++ b/docs/zh-CN/spec/03-runtime/02-agent-runtime.md @@ -695,7 +695,10 @@ Stop / 运行时销毁。主 Agent 用 `TaskStop` 判断要不要取消;运行 回合打开,等委托完成后再把报告塞回父级。父级收工不会中止它们。 致命的 provider/stream 错误(包括耗尽的 HTTP 429)、父级中止,仍分别保留它们既有的 -`failed` 和 `aborted` 结果。 +`failed` 和 `aborted` 结果。如果助手响应在提供程序输出 token 上限处结束(`stopReason: "length"` +或 `"max_tokens"`),且已经产生报告文本,该委派会以 `failed`、 +`SUBAGENT_OUTPUT_TRUNCATED` 和 `outputTruncated: true` 结算;有界的部分报告会保留在失败说明 +下,供诊断截断原因。后续以正常原因结束的委派回合会清除该标记并可以成功完成。 父级终态错误还会中止残留委托、跳过续跑提示,并把会话恢复为空闲,这样 “继续”不会变成 `AGENT_BUSY`(D352)。 diff --git a/docs/zh-CN/spec/03-runtime/08-error-codes.md b/docs/zh-CN/spec/03-runtime/08-error-codes.md index 15d2f9385..bef67e304 100644 --- a/docs/zh-CN/spec/03-runtime/08-error-codes.md +++ b/docs/zh-CN/spec/03-runtime/08-error-codes.md @@ -114,6 +114,7 @@ stdio 与 Tokio 的动态阻塞池隔离,因此后一种情况 | `SUBAGENT_IDLE_TIMEOUT` | 不 | 已撤回(D328):空闲看门狗不再武装;代码仅为已存储结果保留 | | `SUBAGENT_DURATION_TIMEOUT` | 不 | 已撤回(D328):时长看门狗不再武装;代码仅为已存储结果保留 | | `SUBAGENT_CONTEXT_OVERFLOW` | 不 | 委派自身的模型上下文超出其安全预算,自动的回合边界压缩与仅保留任务简报和最近消息的降级重试都没能把它带回限制以内;该失败给出可执行的恢复方式,而不是提供商的溢出文本 | +| `SUBAGENT_OUTPUT_TRUNCATED` | 不 | 委派报告在模型输出 token 上限处结束;保留部分报告用于诊断,但该运行报告为失败,不会被呈现为已完成的委派 | ### 3. 3 工作空间/工具/权限 diff --git a/docs/zh-CN/spec/06-delivery/04-e2e-test-plan.md b/docs/zh-CN/spec/06-delivery/04-e2e-test-plan.md index 2faf3f6fa..41dbfba07 100644 --- a/docs/zh-CN/spec/06-delivery/04-e2e-test-plan.md +++ b/docs/zh-CN/spec/06-delivery/04-e2e-test-plan.md @@ -4539,6 +4539,22 @@ eleven-tool-round desktop paths are verified by - **里程碑**:M6+ - **状态**:草稿。必需套件:`test:e2e`、`test:e2e:subagents`。 +#### E2E-SUBAGENT-output-token-limit-is-a-visible-failure:输出 token 上限必须可见地失败 + +- **先决条件**:Agent 会话使用确定性的本地提供程序;提供程序在产生非空的助手文本后,以 + `stopReason: "length"` 或 `"max_tokens"` 结束。委派有有效报告且没有待处理的工具调用。 +- **步骤**:1)委派任务并让提供程序在输出 token 上限处结束。2)读取 Task 结果、生命周期 + details 和委派卡片。3)用一次正常结束的提供程序响应恢复或重试同一工作。 +- **预期**:第一次运行以 `failed` 结算,而不是 `completed`,并带有 + `SUBAGENT_OUTPUT_TRUNCATED` 和 `outputTruncated: true`。父级收到明确解释,且部分报告保留在 + `Its last output was:` 下面。后续以 `stop` 正常结束的回合会清除标记并以 `completed` 结算; + 部分报告的运行绝不会伪装成已完成报告。 +- **链接规格**:`03-runtime/02-agent-runtime.md` §5f、`03-runtime/08-error-codes.md` §3.2 +- **验收**:C(对话)、H(诊断)、品质 +- **里程碑**:M6+ +- **状态**:由 `packages/agent-runtime/src/subagent.test.ts` 单元覆盖;完整桌面旅程仍需运行: + `test:e2e`、`test:e2e:subagents`。 + #### E2E-SUBAGENT-context-overflow-reports-actionable-failure - **先决条件**:同一个注入的小窗口伪提供商,窗口小到连降级后的上下文也放不下。 diff --git a/packages/agent-runtime/src/subagent.test.ts b/packages/agent-runtime/src/subagent.test.ts index 1b48777aa..63f91dc36 100644 --- a/packages/agent-runtime/src/subagent.test.ts +++ b/packages/agent-runtime/src/subagent.test.ts @@ -367,6 +367,62 @@ describe("SubagentRun reporting", () => { expect(result.error?.code).toBe("SUBAGENT_NO_REPORT"); }); + it("reports truncated output as a failure rather than a clean completion", async () => { + const { run } = createRun(); + run.handleEvent({ + type: "message_end", + message: assistantMessage({ + content: [{ type: "text", text: "Analysis report cut off mid-sentence..." }], + stopReason: "length", + }), + }); + run.agent = { + prompt: vi.fn().mockResolvedValue(undefined), + waitForIdle: vi.fn().mockResolvedValue(undefined), + abort: vi.fn(), + }; + + const result = await (run as unknown as SubagentRun).run(); + + expect(result.status).toBe("failed"); + expect(result.error?.code).toBe("SUBAGENT_OUTPUT_TRUNCATED"); + expect(result.outputTruncated).toBe(true); + expect(result.report).toContain("The explorer subagent failed"); + expect(result.report).toContain( + "The subagent response exceeded the model's output token limit and was truncated.", + ); + expect(result.report).toContain("Analysis report cut off mid-sentence..."); + }); + + it("clears truncated output state if a subsequent turn finishes cleanly", async () => { + const { run } = createRun(); + run.handleEvent({ + type: "message_end", + message: assistantMessage({ + content: [{ type: "text", text: "Initial attempt" }], + stopReason: "length", + }), + }); + run.handleEvent({ + type: "message_end", + message: assistantMessage({ + content: [{ type: "text", text: "Complete final report." }], + stopReason: "stop", + }), + }); + run.agent = { + prompt: vi.fn().mockResolvedValue(undefined), + waitForIdle: vi.fn().mockResolvedValue(undefined), + abort: vi.fn(), + }; + + const result = await (run as unknown as SubagentRun).run(); + + expect(result.status).toBe("completed"); + expect(result.outputTruncated).toBeUndefined(); + expect(result.report).toBe("Complete final report."); + }); + it("returns aborted without prompting when the parent call is already aborted", async () => { const controller = new AbortController(); controller.abort(); diff --git a/packages/agent-runtime/src/subagent.ts b/packages/agent-runtime/src/subagent.ts index cfe8c5eb9..759b3cd86 100644 --- a/packages/agent-runtime/src/subagent.ts +++ b/packages/agent-runtime/src/subagent.ts @@ -110,6 +110,8 @@ export type SubagentRunResult = { contextCompactions?: number; /** True when the run had to discard working history without a summary. */ contextDegraded?: boolean; + /** True when the delegate response hit the model's output token limit. */ + outputTruncated?: boolean; error?: { code: string; message: string }; }; @@ -217,6 +219,7 @@ export class SubagentRun { private readonly opts: SubagentRunOptions; private currentAssistant?: UiMessage; private lastReportText = ""; + private lastReportTruncated = false; private turns = 0; private toolCalls = 0; private usage?: MessageUsage; @@ -375,6 +378,13 @@ export class SubagentRun { message: "The subagent finished without writing a report.", }); } + if (this.lastReportTruncated) { + return this.result("failed", this.lastReportText, { + code: "SUBAGENT_OUTPUT_TRUNCATED", + message: + "The subagent response exceeded the model's output token limit and was truncated.", + }); + } return this.result("completed", this.lastReportText); } @@ -648,6 +658,7 @@ export class SubagentRun { ...(this.modelFailures.length ? { modelFailures: [...this.modelFailures] } : {}), ...(this.contextCompactions > 0 ? { contextCompactions: this.contextCompactions } : {}), ...(this.contextDegraded ? { contextDegraded: true } : {}), + ...(this.lastReportTruncated ? { outputTruncated: true } : {}), ...(error ? { error } : {}), }; } @@ -826,6 +837,8 @@ export class SubagentRun { // must not clear the text an earlier turn already produced. if (content.hasText && content.text.trim() && !failed) { this.lastReportText = content.text; + this.lastReportTruncated = + stopReason === "length" || stopReason === "max_tokens"; } if (retryAttempt !== undefined) { this.currentAssistant = { diff --git a/packages/shared/src/errors.test.ts b/packages/shared/src/errors.test.ts index ebabdf3af..4b78b3bc5 100644 --- a/packages/shared/src/errors.test.ts +++ b/packages/shared/src/errors.test.ts @@ -35,6 +35,7 @@ describe("result helpers", () => { const newlyLiveCodes = [ "CONTEXT_COMPACTION_FAILED", "EMPTY_MODEL_RESPONSE", + "SUBAGENT_OUTPUT_TRUNCATED", "EDIT_TAG_REQUIRED", "EDIT_TAG_MISMATCH", "EDIT_TAG_UNKNOWN", diff --git a/packages/shared/src/errors.ts b/packages/shared/src/errors.ts index 884ccc181..e8654ff11 100644 --- a/packages/shared/src/errors.ts +++ b/packages/shared/src/errors.ts @@ -107,6 +107,7 @@ export const ErrorCodes = { * delegate's model, or how much it reads at once has to change. */ SUBAGENT_CONTEXT_OVERFLOW: "SUBAGENT_CONTEXT_OVERFLOW", + SUBAGENT_OUTPUT_TRUNCATED: "SUBAGENT_OUTPUT_TRUNCATED", WORKSPACE_REQUIRED: "WORKSPACE_REQUIRED", PATH_OUTSIDE_WORKSPACE: "PATH_OUTSIDE_WORKSPACE", TOOL_NOT_FOUND: "TOOL_NOT_FOUND",