fix(storage): preserve shared mount parent during lease cleanup - #2008
Merged
Merged
Conversation
Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com>
5 tasks done
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The focused change addresses the reported race, the regression covers its failure mechanism, and no unresolved issues were identified.
Review effort: Balanced
Findings: None
What changed in this PR
This PR fixes a native-mount lease cleanup race by keeping the shared resource directory available to concurrent lease creators.
Changes:
- Stop removing the shared parent when a lease is retired.
- Add a directory-handle regression test and a changelog fragment.
| File | Description |
|---|---|
crates/astrid-kernel/src/storage_mount/tests.rs |
Includes the new cleanup tests. |
crates/astrid-kernel/src/storage_mount/cleanup_tests.rs |
Tests creation through an open parent handle after cleanup. |
crates/astrid-kernel/src/storage_mount.rs |
Preserves the shared parent during lease cleanup. |
changes/2006.fixed.md |
Records the fix. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
joshuajbouw
added a commit
that referenced
this pull request
Sep 29, 2026
## Linked Issue Closes #2001 #1996 is merged. The retention-only commits are rebased onto main `ed2df49b`, preserving the Wasmtime 48.0.3 security update. Current head: `896c6e4c3d9f5fb305278a3a408250a7908d7204`. The integration adds one Windows Clippy fix: directory syncing and its stored path are compiled only on Unix; Windows retains the existing write-through rename behavior. No global lint suppression. ## Summary Today, pruning deletes entries permanently. That holds whether an operator runs it or the global cap triggers it (1,000,000 entries / 1 GiB), and only a chain's latest prune receipt is kept. So history no external party has seen can be erased, and a verifier cannot connect retained entries to history it anchored several prunes earlier. This PR changes that: - An anchoring service records, per chain, how far it has certified the log: an anchored watermark. - Retention never deletes past a watermark. When the cap is reached and nothing can be pruned, the log keeps the entries and reports degraded instead of deleting them. - Every prune receipt is kept and can be exported. - Operators can require a watermark before any prune, and can archive pruned segments. Installs that do not anchor keep today's bounded retention. ## Changes - **`astrid-audit`: anchored watermarks.** - `AuditLog::mark_anchored` records that a chain was certified through a position. The position counts entries from the chain's genesis, pruned ones included (`omitted_total + count`). The mark carries the head hash at that position and up to 4,096 bytes of opaque evidence. - The claim is checked against the chain: the entry at `position - 1`, or, at the pruned boundary, the latest receipt's terminal hash, must hash to the given head. - Verification resumes from the previous watermark's stored cursor, so a mark costs only the entries appended since. - Watermarks never decrease. Re-marking the same position and hash is a no-op. - Marks live in their own namespace (`audit:anchor_marks`) and are replaced by compare-and-swap. Chain metadata, which append-intent recovery compares byte for byte, is untouched. - **`astrid-audit`: prunes never pass a watermark.** - Each prune is checked against the chain's pruned total after its retention scan. The check runs again, under the durable append lock, when the deletion plan is created. - Marks are installed under that lock only while no plan exists and the receipt is unchanged since verification. This closes the window in which a chain's first mark could land between a prune's check and its plan. - A refused prune deletes nothing and writes no receipt. - `prune_oldest` (used by `audit prune` and the cap) takes the oldest sealed segment, in seal order, whose removal is allowed. It searches the whole seal-order index before refusing. - **`astrid-audit`: the cap holds instead of deleting.** When an append reaches the cap and every candidate holds unanchored history: - a durable retention hold is set in the global metadata; - appends are admitted over the cap; - the global state reports degraded with the reason, and an error is logged once. The hold clears when a watermark advances, a prune finishes, a segment is sealed or the caps change. - **`astrid-audit`: every receipt is kept.** Receipts are recorded under `audit:receipt_history/<chain>/<generation>` before installation. A chain pruned before this change keeps its latest receipt. `AuditLog::prune_receipts` pages them. - **`astrid-audit`: archive, then drop.** With an `AuditArchiver` set, a prune writes the entries it removes to an archive and commits it before the deletion plan is created. The archive must match the receipt's count, digest and terminal hash, or the prune is abandoned. Prunes are serialized from segment selection to plan installation, so two prunes of one chain cannot archive the same generation. - **`astrid-core`, `astrid-kernel`: admin methods.** - `audit.anchor_mark` (capability `audit:anchor`) accepts one certification's evidence and up to 4,096 chains. Each chain is reported as advanced, unchanged or rejected with the reason. Malformed evidence, an empty or oversized list, or a duplicate chain rejects the whole request. - The kernel checks the evidence's shape, not the certification itself. - Successful marks skip the generic success row, as `audit.heads` and `audit.export` do. Rejections and denials are recorded. - `audit.anchor_status` (capability `audit:heads`) reports each chain's count, pruned total, watermark, its time and evidence, whether anchoring is required before prune, and the retention hold. - `audit.export` gains `receipts_from`, which pages prune receipts with their hashes and signing bytes. - `audit.stats` reports the hold. - The file archiver does its filesystem work on the blocking thread pool. - **`astrid-config`: `[audit.retention]`, operator-only.** - `require_anchor` (default `false`) makes every prune require a watermark. - `archive_dir` must be an absolute path. It writes each pruned segment as `<session>/<system|principal.alias>/<generation>.jsonl`: the signed receipt, then the removed entries. The file is written under a temporary name, synced and renamed, in directories private to the daemon's user. - A workspace layer cannot override either key. - `AuditConfig` moves to its own module. - **`astrid-cli`:** - `astrid audit anchor-status` shows per-chain watermarks and lag, and exits 2 while the hold is set. - `audit prune` explains a refusal. - `audit stats` shows the hold. - `audit export --receipts-from` prints prune receipts. - Scenario entries. - Changelog fragment. New storage fields are optional and absent by default, so existing metadata re-encodes to its stored bytes. ## Verification Current integration validation: audit library 99 passed / 2 ignored; config library 138 passed; kernel retention 2 passed; kernel-library Clippy with warnings denied, formatting and diff checks passed on macOS. All 33 checks on 896c6e4 now pass, including Windows and both MUSL targets. #2008 is merged. The exact combined tree with main bdc1fe8 (tree 6c15a1d) also passes the full kernel library suite: 605 passed, zero failed or ignored. Prior author validation of the retention implementation follows: - **Test suites:** `astrid-audit` 99 (+2 doc), `astrid-core` 390, `astrid-config` 138, `astrid-uplink` 58 and CLI 818 all pass. The kernel lib passes 602 of 603 with `--test-threads=6` and umask 0022. The remaining failure needs a copy-on-write workspace backend and also fails on the base commit. With full parallelism on a loaded host, some kernel-router tests exceed their 2-second admin-response timeout; #1996's branch shows the same under the same load. - **Unit and handler tests:** - watermark verification and monotonicity; - positions across prunes and at the pruned boundary; - persistence across reopen; - a first mark recorded while a prune is held between its check and its plan: the prune is refused with nothing deleted; - a mark refused while a prune is pending; - prune refusal past a watermark; - the require-anchor switch; - oldest-eligible segment selection across refused chains; - the cap hold, which keeps every append and clears on anchoring; - cap pruning unchanged for chains without watermarks; - receipt history order and backfill; - archive contents and permissions, and abort on archive failure; - two concurrent prunes of one chain archiving one at a time; - per-chain `anchor_mark` outcomes and whole-request rejections; - receipt pages in export; - no row written for a successful mark, while rejections and denials write one; - config restriction and validation. - **Lint:** `cargo fmt` and `cargo clippy --all-features --all-targets -- -D warnings` on the touched crates. ## AI / Tool Assistance Assisted-by: Claude Code:claude-opus-5-5. Covers design, implementation and tests in all touched crates, and this description. Assisted-by: Codex CLI. Review passes over the diff. #1996 and #2008 are merged; all current-head CI checks pass. ## Checklist - [x] Linked to an issue - [x] Changelog fragment added under `changes/{issue}.{kind}.md` (docs/CI-only may skip; release PRs roll fragments into the version section instead of adding one) - [x] I understand every change in this PR and can explain its design, risks, and validation. - [x] I reviewed and tested any meaningful tool-generated output included in this PR. - [x] Every non-bot, non-merge commit has a matching `Signed-off-by` trailer. --------- Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com> Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com> Co-authored-by: Joshua J. Bouw <jjb@unicity-labs.com>
5 tasks done
joshuajbouw
added a commit
that referenced
this pull request
Sep 29, 2026
## Linked Issue Closes #1999 Part of #1997. #1996, #2002, and the mount cleanup fix #2008 are merged. Current head `c92562e65087fe47e313e4c029781e89a69927d4` is based on main `26a500a6`. Remaining merge order: this PR, #2004 (recording coverage), #2005 (entry format v2). The integration preserves Pavel's ordered host-audit implementation. The overlapping configuration is resolved in the existing `astrid-config::audit` module, retaining both retention validation and fail-closed host-call settings. Kernel startup loads both settings from the same audit configuration. This PR replaces the host-audit writer. #2004 changes the old writer's fold key to include the capsule. Whichever of the two lands second is rebased onto the other, and the capsule then becomes part of the run key. ## Summary The host-audit sink could lose calls and reorder them without a trace: - It coalesced allowed calls per class over `host_coalesce_ms`, and appended `repeats=N` to the outcome details. The entry signature covers the outcome only as a success bit, so that count was not signed. - It drained pending work per key, so a chain's entries did not follow call order. - A call that met the full queue (`host_queue_capacity`, default 4,096 keys) was dropped, and only a counter moved. - A failed batch was dropped, and calls still queued when the daemon died vanished. As a result, even an externally anchored log could not show that it was complete. This PR records host calls in call order, and makes every loss visible as a signed entry on the chain. It also lets an operator make chosen host-call classes fail closed. ## Changes - **`astrid-audit`: signed host-call records.** These actions are covered by the entry signature: - `HostCallRun`: consecutive calls of one principal. It holds the call count, the first and last call times, a tally per class and outcome, and a fold over every call. - `HostCallLoss`: calls that were accepted but not recorded individually. - `HostCallGap`: an earlier run of the lane stopped without draining. - `HostCallAdmitted`: the write-ahead entry of a fail-closed call. `astrid_audit::host_call` specifies the per-call digest and the fold, so a verifier can recompute them without the kernel. Known-answer tests pin them. - **`astrid-kernel`: ordered lane** (`audit_sink/lane.rs`, `writer.rs`). - Each principal chain has a FIFO of pending slots, so FIFO order is call order. - The writer takes slots from the front and appends each batch atomically with `append_batch_with_principal`. - A failed batch is retried as is, with backoff up to 5 s, before anything behind it is taken. - Signing happens at the durable commit, because only the commit fixes a chain position. The kernel router appends to the same chains directly. - **`astrid-kernel`: lossless coalescing.** - Consecutive allowed and failed calls share one run, and consecutive identical denials share one run. - A one-call run is written as the plain action, as before. - **`astrid-kernel`: loss and gap entries.** - Queue capacity bounds slots, not calls. Overflow folds into one loss slot per chain, which is written as `HostCallLoss` at its place in the chain. - A lane marker in `system:control:audit-lane` records the run and the chains it wrote. After an unclean stop, the next start writes `HostCallGap` into each of those chains and into the session's system chain. - `astrid-storage`: the state-owner resolver admits `system:control:audit-lane` as a system control projection, like `audit`, `invites` and `pair-tokens`. - An unreadable marker, or a marker store that fails, also produces a gap entry. - If the writer thread dies, the lane closes and health reports it. - **`audit.host_fail_closed`** (`astrid-config`, `astrid-capsule`, `astrid-kernel`). - It lists host-call classes whose effect runs only after a write-ahead entry is durable: `file_read`, `file_write`, `file_delete`, `net_connect`, `net_bind`, `process_spawn`. The default is empty. - `HostAuditSink::admit()` is called by the fs, net and process host functions after the security gate and before the effect. - A refusal fails the call with `unknown("audit unavailable")` and is recorded as a denial. - A workspace layer can add classes but not remove them. - **Health:** `accepted` and `persisted` now count calls, and `lost`, `gaps_recorded` and `dropped_after_shutdown` are new. The admin `AuditHealth` wire type is unchanged. - Changelog fragment. **Consequences.** - A denial between allowed calls closes the run. - Order is guaranteed among a chain's host-call entries. Admin rows are still appended when they run. - Calls reported after shutdown are counted in health, not written. **Overhead** (release build, in-memory log, per producer thread; `audit_sink::tests::bench`, ignored by default): | Workload | Before | After | |---|---|---| | 1 thread, 100k mixed allowed calls | 1.61–1.65 µs/call | 0.84–0.92 µs/call | | 8 threads × 8 principals, 200k calls | 15.1–17.9 µs/call | 8.9–12.8 µs/call | | 4 threads, reads with 2% denials | 6.7–7.8 µs/call | 4.2–4.8 µs/call | | 20k distinct denials, 1 thread | 1.35–1.58 µs/call; 15,904 calls dropped | 1.82–1.91 µs/call; none dropped, overflow in 5–7 loss entries | ## Verification Current integrated tree: full kernel library **628 passed, 1 ignored**; capsule audit tests **45 passed**. Earlier validation of the same recorder/retention integration: audit 106 passed / 2 ignored, config 140 passed, kernel audit filter 51 passed / 1 ignored; audit/config/kernel library Clippy passed with warnings denied. The published tree is byte-identical to the tested combined tree; CI must pass on this new head before merge. Prior author verification (before this integration): - **Test suites:** `astrid-audit` 74, `astrid-capsule` 748, `astrid-config` 139 and `astrid-storage` 864 all pass. The kernel lib passes 610 of 611 with `--test-threads=6` and umask 0022. The remaining failure needs a copy-on-write workspace backend and also fails on `main`. - **Unit and behaviour tests:** - host-call digest and fold known-answer values; every field changes the digest; - a dropped, reordered or altered call fails `matches_calls`; - changing a signed count breaks the signature; - call order across 6 producers, 3 principals, 8-entry batches and a 10 ms window (the same probe fails on `main`); - a 40-call run whose fold, count and time range recompute from the calls; - queue overflow written as one loss entry that commits to the lost calls; - gap entries after an abandoned lane run, none after a clean restart, and gap duties carried across a second unclean stop; - a corrupt marker and a failing marker store, including the unavailable-marker gap surviving a later marker write; - a system-only gap recorded at start with no host call, and retried while idle after a refused append; - the marker namespace written and read back through a runtime principal store; - a batch the log refuses, retried and written in order once the cap is raised; - a fail-closed write durable before `admit()` returns; - refusal within the first failed batch, and after shutdown; - config validation and workspace restriction; - `connect-tcp` and `bind-tcp` refusing without a connection attempt. - **Lint:** `cargo fmt` and `cargo clippy --all-features --all-targets -- -D warnings` on the touched crates. ## AI / Tool Assistance Assisted-by: Claude Code:claude-opus-5-5. Covers design, implementation and tests in all touched crates, and this description. Assisted-by: Codex CLI. Review passes over the diff. Current integration and verification assisted by Codex. Fresh CI pending; no claim of merge readiness until those checks complete. ## Checklist - [x] Linked to an issue - [x] Changelog fragment added under `changes/{issue}.{kind}.md` (docs/CI-only may skip; release PRs roll fragments into the version section instead of adding one) - [x] I understand every change in this PR and can explain its design, risks, and validation. - [x] I reviewed and tested any meaningful tool-generated output included in this PR. - [x] Every non-bot, non-merge commit has a matching `Signed-off-by` trailer. Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com> Co-authored-by: Joshua J. Bouw <jjb@unicity-labs.com>
5 tasks done
joshuajbouw
added a commit
that referenced
this pull request
Sep 29, 2026
… changes and capsule loads (#2004) ## Linked Issue Closes #1998 Part of #1997. #1996, #2002, #2003, and #2008 are merged. Current head `c0f4052305da424d9625c9f96332448667a9259b` integrates this PR with main `e63f65cc`. Remaining order: this PR, then #2005 (entry format v2). The integration retains #2003's ordered, loss-accounted writer and fail-closed admission. Capsule identity is part of the run key, so alternating capsules cannot fold together. HTTP/tool/approval coverage records remain individual queue entries; overflow still records a payload-bound loss summary. Existing unattributed call digests keep their exact v1 encoding; attributed and new coverage calls use the documented `astrid.audit.host-call.v2` digest domain. This is a call-fold encoding distinction, not #2005's entry-format change. Ordering boundary: queued records retain per-principal FIFO order. Synchronous HTTP/approval commits still use the existing durable append path, like direct administrative audit writes; no new global ordering guarantee between that path and queued records is claimed. ## Summary The audit log did not record the events that determine what an agent did: - LLM and other HTTP exchanges appeared only in `tracing::debug!`, so an agent turn, including its model request, added no entry; - capsule tool calls and approval decisions had no producer; - capability grants and revocations were recorded as `AdminRequest` rows at authorization time, before the change was applied and without its result, and grant-on-use was not recorded at all; - capsule install and load were not recorded, and install hooks ran without an audit sink; - host-call entries did not name the capsule, and `FileWrite` stored a zero hash. This PR records all of these on the signed log, without changing the v1 entry format. Content never enters the log, only commitments to it. Provider credentials can now be injected by the host, so they never enter guest memory. ## Changes - **`astrid-audit`: entry kinds.** - `HttpRequest` (pre-commit) and `HttpResponse` (completion). - `CapabilityChanged`, `CapsuleInstalled` and `CapsuleLoaded`. - A `CapsuleActor` (capsule id and wasm hash) on file, network, process, tool-call and approval entries. - Linking fields on `CapsuleToolCall` and the approval entries. New fields are optional and omitted when unset, and new variants are appended after the existing ones. Entries written before this change re-serialize to the same signed bytes. - **HTTP exchanges** (`astrid-capsule` HTTP host, `astrid-kernel`). - Each wire request, each redirect hop included, gets a durable `HttpRequest` pre-commit before it is sent. The pre-commit holds the method, host, port, BLAKE3 commitments to the path, headers and body, the redirect hop, and a per-principal sequence number scoped to the kernel run. - An `HttpResponse` completion links to its pre-commit. It carries the status, provider request ids, and a hash of the response body as it was read (incrementally for streams). - Completions are appended durably. - Refusals by the scheme check, egress, the security gate or the SSRF airlock are recorded as denied requests and take a sequence number. - Credential headers, credential query parameters and every secret value the capsule received are redacted before hashing. The redaction rules are documented so a verifier can recompute the commitments. - Recording is best-effort, like the rest of the audit log since 0.9.0. If the pre-commit cannot be written, the request is still sent, a security event is logged, and the missing entry shows as a gap in the run's sequence numbers. Approval prompts behave the same way. Making these fail closed is an operator policy; it would fit as an `http` class of `audit.host_fail_closed` from #2003 once both have landed. - **Host-side credentials.** - A header value can name a manifest-declared secret as `{{secret:NAME}}`, and the host substitutes it on the wire. - Undeclared names are refused and recorded. Placeholders are stripped on a cross-origin redirect. - Commitments cover the placeholder form, and the entry lists injected secret names, never values. - `docs/models.md` describes how a provider capsule adopts this. - **Tool calls and approvals.** - One `CapsuleToolCall` entry per `tool.v1.execute.<tool>` invocation, with the call id, argument and result hashes, and the capsule. - Approval prompts are committed before they are published. Every decision is linked to its prompt, with how it was reached: user, session allowance, remembered consent, timeout and so on. - Grant-on-use prompts are committed by the dispatcher before `GrantRequired` is published. An approval is committed before `GrantResult` is published, so the chain shows it before the grant it caused. - **Capability changes.** Applied changes are recorded on the affected principal's chain after they are saved: token mint and revoke, caps grant and revoke (only the patterns actually added), group and capsule membership changes, distro self-grants, and grant-on-use. - **Code identity.** - `CapsuleInstalled` binds the wasm and exact `Capsule.toml` hashes. - `CapsuleLoaded` (on load or replace) binds the verified wasm hash, the manifest hash and the engine profile. - Install and upgrade hooks run by the daemon get the signed sink, attributed to the hook's component. - **Attribution.** - Host-call entries name the capsule and wasm hash, stamped by the host from the verified component; a guest cannot choose them. - `write-file` reports the hash of the bytes written. - Changelog fragment. ## Verification Current signed integration head: **111 audit tests passed / 2 ignored; 777 capsule tests passed; 643 kernel tests passed / 1 ignored** (1,531 passed in total). Commands: `cargo test --locked -p astrid-audit -p astrid-capsule -p astrid-kernel --lib -- --quiet`; `cargo check --locked --workspace`; `cargo clippy --locked -p astrid-audit -p astrid-capsule -p astrid-kernel --lib --tests -- -D warnings`; formatting and diff checks. All passed locally on macOS. Added integration regressions bind capsule/wasm identity in folded digests and prove coverage events remain individual while overflow binds the original payload. The attribution test now asserts FIFO order across alternating capsules instead of expecting the old unordered folding behavior. Fresh platform CI is still required. Prior author verification before integration: - **Test suites:** `astrid-audit` 71, `astrid-capsule` 777, `astrid-capsule-install` 106 and CLI 818 all pass. The kernel lib passes 601 of 602 with `--test-threads=6` and umask 0022. The remaining failure needs a copy-on-write workspace backend and also fails on `main`. - **Unit and integration tests:** - legacy JSON encodings pinned, and a legacy entry re-verified after a round trip; - pre-commits resolve before a loopback server sees the request; - buffered and streamed responses hashed and linked, including streams closed early, transport failures and gate denials; - redaction: short secrets, many secrets, overlapping secrets; - credential injection end to end: the server receives the secret, the guest never reads it, and the commitment covers the placeholder; - placeholder rules: undeclared, malformed and unterminated placeholders refused, also after a blank secret, and an omitted header names no secret; - tool results captured through the IPC host function, with error, missing, mismatched and cancelled cases; - approval prompts committed before they are visible on the bus; - capability grant, revoke and membership entries, with none for a no-op re-grant, and a grant-on-use approval ahead of its grant; - a daemon install producing `CapsuleInstalled` and `CapsuleLoaded` with the expected hashes, and an upgrade recorded as load then replace. - **Lint:** `cargo fmt` and `cargo clippy --all-features --all-targets -- -D warnings` on the touched crates. `cargo check --target wasm32-unknown-unknown` of `astrid-audit`, `astrid-capsule` and `astrid-kernel`. ## AI / Tool Assistance Assisted-by: Claude Code:claude-opus-5-5. Covers design, implementation and tests in all touched crates, and this description. Assisted-by: Codex CLI. Review passes over the diff. Integration onto the landed recorder and associated regression tests assisted by Codex. The earlier source review does not constitute independent review of these integration edits. ## Checklist - [x] Linked to an issue - [x] Changelog fragment added under `changes/{issue}.{kind}.md` (docs/CI-only may skip; release PRs roll fragments into the version section instead of adding one) - [x] I understand every change in this PR and can explain its design, risks, and validation. - [x] I reviewed and tested any meaningful tool-generated output included in this PR. - [x] Every non-bot, non-merge commit has a matching `Signed-off-by` trailer. --------- Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com> Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com> Co-authored-by: Joshua J. Bouw <jjb@unicity-labs.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Linked Issue
Closes #2006
Summary
Keep the private shared mount container alive when retiring an individual lease. Concurrent lease creation can already hold a directory handle to that container; unlinking it makes the next mkdirat fail with ENOENT.
This is a pre-existing lifecycle bug investigated after the Ubuntu failure on #2002, not an audit-retention change. The deterministic regression reproduces the same error class; the original CI interleaving was not captured.
Changes
Verification
Test Plan
Run the regression and platform CI. Confirm cleanup still removes lease-owned resources and preserves the parent needed by another creator. The retained empty parent is intentional, not a leaked lease.
AI / Tool Assistance
Assisted-by: Codex. Traced the cleanup/creation race, wrote the focused fix and regression, and executed before/after validation.
Checklist