How should stateful cleanup failures be surfaced? #1036
PierrunoYT
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Policy question
Several ownership boundaries currently represent cleanup as
func()or discard terminal errors, so callers cannot distinguish best-effort hygiene from teardown that may leave locks, processes, policy grants, or transports behind.Examples at audited
mainrevision1b5db1765672820caac1684b168c9898b5ba3593include:func()and both errors are ignored (internal/mcp/permissions.go,lockStateFile);func(), so adapter/PTY teardown cannot contribute to the process result (internal/execution/process_manager.go);WorkerHandle.Kill()errors (internal/daemon/pool.go);func()callbacks with no result (internal/agent/loop.go).Routine response-body closes and temporary-file removal should not automatically fail successful operations. The question is which cleanup failures change durable or externally observable state and therefore must be returned, joined, or recorded.
Audit context: finding
REL-01. No single user-visible failure was reproduced, so this is a design/ownership discussion rather than one bundled bug report.Proposed classification
Questions
func() error, and which should remain best-effort?errors.Join, or preserve a primary error plus diagnostics?Acceptance properties for later scoped work
errors.Is/errors.As;Full audit detail: https://github.com/PierrunoYT/zero/blob/audit/codebase-audit-2026-09-08/docs/audit/CONCURRENCY_AUDIT.md#rel-01--stateful-cleanup-errors-are-inconsistently-observable
All reactions