feat(audit): retention that never deletes unanchored history - #2002
Conversation
joshuajbouw
left a comment
There was a problem hiding this comment.
Reviewed 91f2742. No blocking findings. The retention policy stays operator-controlled, watermark checks are repeated under the storage lock before a prune plan is accepted, and archive failures stop deletion. This keeps external anchoring policy outside Astrid while providing the runtime mechanisms it needs.
Independently ran 639 tests across audit, config, core, and the focused kernel retention/anchor handlers (2 ignored), plus focused Clippy with warnings denied. All passed. The checks cover stale watermarks, receipt history, archive contents and permissions, and refusing deletion when archiving fails.
This is a source review of the draft, intended to follow #1996. GitHub currently reports no CI checks for this branch; these local results are not a claim of CI completion or live external-anchoring certification.
|
#1996 is now merged as ed2df49, including the Wasmtime 48.0.3 security fix. Please rebase this branch onto main as planned: it currently conflicts in the shared heads/export and admin-handler files, so I cannot merge it yet. The current CI also has two separate issues beyond the now-fixed dependency audit:
The retention review remains useful. After the rebase, I will review the integration delta and refreshed results rather than repeat the unchanged review. Keeping the agreed merge order: #2002, #2003, #2004, #2005. |
|
Follow-up on the Ubuntu ENOENT: storage_mount.rs is byte-identical between this head and the now-landed #1996 head. There is a plausible existing race: private_mount_resource_path ensures the shared /tmp/astrid-mounts-UID directory, then issue_lease creates its UUID child separately; meanwhile cleanup_resource removes that shared parent after removing another lease. A cleanup between those two creation steps matches the observed error. This is source-level diagnosis, not yet a reproduced interleaving, so I would not attribute it to retention or simply label it infrastructure. The immediate integration step remains the planned rebase plus the Windows sync_directory lint fix. Please preserve the shared heads/export changes and the Wasmtime update from main. |
|
Confirmed the shared mount-parent lifetime defect with a deterministic regression on main: hold the shared parent directory open, retire the previous lease via the real cleanup_resource, then mkdirat for the arriving lease fails with ENOENT. This reproduces the same error class as Ubuntu CI, although the original CI interleaving was not recorded. I am preparing the narrow fix separately under #2006 so it does not add unrelated source changes to your retention PR. |
|
The mount-parent fix is now #2008. The new regression fails with ENOENT before the fix and passes afterward; all 598 kernel library tests and focused Clippy pass locally. Platform CI is still pending. This addresses the reproduced cleanup race separately from retention. Your rebase and Windows archive-helper lint correction are still needed. |
Pruning deleted entries permanently, whether an operator asked for it or the global cap (1,000,000 entries / 1 GiB) triggered it on an append, and only the latest prune receipt of a chain was kept. History that no external party had seen could be erased, and a verifier could not connect retained entries to history it anchored before several prunes. Anchored watermarks. AuditLog::mark_anchored records that an external anchoring service certified a chain through a position: entries counted from the chain's genesis, pruned ones included (omitted_total + count in audit.heads), with the head hash at that position and opaque evidence of at most 4,096 bytes. The claim is checked against the chain: the entry at position - 1 must hash to the given head hash, or, at the pruned boundary, the latest receipt's terminal hash must. Verification reads onward from the previous watermark's stored cursor, so repeated marks cost only the entries appended since. A watermark never decreases: a lower position is rejected, and the same position and hash succeed without change. Marks live in their own namespace (audit:anchor_marks) and are replaced by compare-and-swap, so chain metadata, which append-intent recovery compares byte for byte, is untouched. A prune that runs while a mark is verified is detected (pending plan or changed receipt) and the mark is rejected for retry. Prunes never pass a watermark. Every prune is checked after its retention scan against the chain's pruned total, read after the prior receipt its own receipt links to; a prune that finishes in between breaks that link and the new plan is refused, so the check never sees a stale total. The check is repeated when the deletion plan is created, under the durable append lock, and marks are installed under that lock only while no plan exists and the receipt is unchanged since verification. Without that, a chain's first mark recorded between a prune's check and its plan let the prune delete past the new watermark. A refused prune deletes nothing and writes no receipt. prune_oldest, used by `audit prune` and by the cap, now takes the oldest sealed segment in seal order whose removal is allowed, skipping chains whose oldest segment is not (their later segments are not either), and returns the refusal for the oldest when none is. Chains without a watermark are pruned as before unless AuditLog::set_require_anchor_before_prune(true). The cap alarms instead of deleting. When an append reaches the cap and every candidate would remove unanchored history, the log sets a durable retention hold in the global metadata: appends are admitted over the cap, the global state reports degraded with the reason, and an error is logged once. The hold clears when a watermark advances, a prune finishes, a new segment is sealed or the caps change, and the next append at the cap tries to prune again. The hold is written under the durable append lock after recovering any interrupted append, as appends require. Every receipt is kept. Each receipt is recorded under audit:receipt_history/<chain>/<generation> before it is installed; the receipt it replaces is recorded too, so chains pruned before this change keep their latest one. AuditLog::prune_receipts pages them by generation. Archive-then-drop. With an AuditArchiver set, a prune streams the entries it removes to an archive and commits it before the plan that deletes them is created. The archived entries must match the receipt's count, digest and terminal hash, or the prune is abandoned with nothing deleted. New storage fields are optional and absent by default, so existing metadata re-encodes to its stored bytes. Tests: watermark verification and monotonicity, positions across prunes and at the pruned boundary, per-chain watermarks, persistence across a store reopen, a first mark recorded while a prune is held between its check and its plan (the prune is refused with nothing deleted; it deletes past the mark without the plan-time check), a mark refused while a prune is pending, prune refusal past a watermark with nothing deleted, the require-anchor switch, oldest-eligible segment selection, the cap hold keeping every append and clearing on anchoring, unchanged cap pruning without watermarks, receipt history order and backfill, and archive contents and abort on archive failure. The existing audit tests pass unchanged apart from two test-only AuditLog constructions. Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com>
The anchoring service had no way to tell the kernel what it anchored, so retention could not respect it, and neither operators nor verifiers could see how far anchoring lagged or read earlier prune receipts. audit.anchor_mark (AdminRequestKind::AuditAnchorMark, capability audit:anchor) takes the evidence of one certification (network, checkpoint digest, link position, and optionally the certifying round and seal digest) and up to 4,096 chains, each with a position and head hash. The evidence's shape is checked and it is stored with each chain's watermark; the kernel does not verify the certification itself. Each chain is checked and recorded independently and reported as advanced, unchanged or rejected with the reason. Malformed evidence, an empty or oversized list, or a chain listed twice rejects the whole request. Successful marks skip the generic success row, as audit.heads and audit.export do: each row would be a new unanchored entry. A rejection is recorded as a failure row naming the rejected chains; the row stores the chain count instead of the full list. audit.anchor_status (capability audit:heads, which reveals the same per-chain positions and hashes) reports every chain's count, pruned total, watermark, its time and evidence, whether anchoring is required before prune, and the retention hold. It is bounded to 4,096 chains like the heads snapshot. audit.export takes receipts_from and then returns the chain's prune receipts from that generation on, 32 per page, each with its hash and signing bytes, inside the same settled-prune window as the entries. audit.stats reports the retention hold. A prune refused by a watermark is reported as "audit prune refused: <reason>". Tests cover per-chain outcomes, whole-request rejections, the refusal through audit.prune with a missing and a partial watermark, the receipt history in export pages, and that successful marks write no row while rejections and denials do. The kernel_router suite passes. Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com>
Whether a chain without a watermark may be pruned, and whether pruned segments are archived first, are operator decisions that belong in the daemon config. [audit.retention] require_anchor (default false) makes every prune require an anchored watermark, so an install that anchors never loses unanchored history to the cap; installs that do not anchor keep bounded retention. [audit.retention] archive_dir, an absolute path, makes every prune first write <session>/<system|principal.alias>/<generation>.jsonl: the signed receipt, then each removed entry as stored. The file is written under a temporary name, synced and renamed into place, its directories are made private to the daemon's user, and a failed archive leaves the chain intact. The table is operator-only: a workspace layer cannot turn off the requirement or move the archive. AuditConfig moves to its own module next to the new table, which also shrinks types.rs. The kernel applies the table when it boots. A config that fails to load falls back to the defaults, as the host-audit settings already did, now with a warning. Tests: the restriction reverts a workspace's table, archive_dir must be absolute, and an archive holds the receipt and every removed entry in chain order with private permissions and no leftover temporary file, while an unwritable archive leaves the chain unpruned. Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com>
`astrid audit anchor-status` lists each chain's head position, anchored watermark, lag and when it was recorded, whether anchoring is required before prune, and the retention hold; it exits 2 while the hold is set, as `audit stats` does when degraded. `astrid audit prune` explains a refusal and points to anchor-status, `audit stats` shows the hold or the last retention error, and `audit export --receipts-from` prints the chain's prune receipts. The command is mapped in cli-scenarios.toml. Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com>
…time#2001 Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com>
The file archiver created directories, wrote every archived entry, flushed and synced the file, and renamed it directly inside its async methods. A prune of a large segment, or a slow disk, then held a runtime worker thread for the whole archive and its fsyncs, stalling unrelated kernel requests on that worker; an automatic prune does this inside an append. Directory creation, file creation and the receipt line now run in one blocking task when the archive opens; each batch of entries is serialized on the async side and written in a blocking task that hands the file back; and flush, sync, rename and the directory sync run in one blocking task at commit. An unfinished archive still removes its temporary file when dropped. The archive tests pass unchanged. Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com>
Choosing the segment to prune stopped after 64 distinct chains or 4,096 index keys and returned the oldest refusal. With 64 chains whose oldest segments hold unanchored history ahead of a prunable one, `audit prune` refused and the global cap set its retention hold, and every later attempt scanned the same prefix: the prunable segment was never reached and storage grew past the cap. The search now reads the whole seal-order index before it refuses, so a refusal means no sealed segment can be pruned. It still skips, without reading them, the later segments of a chain already refused, which the watermark refuses too; a full search costs a page of keys per 256 segments and a few reads per chain, and while the hold is set appends at the cap do not search again until something changes what can be pruned. The selection test now places a refused chain's second segment between the refused and the prunable chains. Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com>
Two prunes of one chain, such as `audit prune` and a prune the global cap starts inside an append, read the same prior receipt and sign receipts of the same generation. Both could archive before either installed its deletion plan. The file archiver names an archive after its generation, so the second archive replaced the first; if the first prune's plan was then installed, it deleted entries that the surviving archive does not hold. Prunes now hold a process-wide lock from choosing what to remove until the deletion plan is installed and finished: `prune_chain` around the scan, archive and plan, and `prune_oldest` from segment selection on, so the chosen segment is still there. The lock is taken after an append's chain lock, which the cap prune runs under, and before the durable append lock; nothing that holds it appends. A test holds the first of two concurrent prunes in its archive and checks that the second does not start archiving until the first has finished; without the lock both archive generation 0. Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com>
Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com>
|
I prepared a ready-to-use integration branch to save the rebase work: astrid-runtime/astrid branch integrate/audit-retention-2002, head 896c6e4. All eight retention-only commits apply cleanly onto main ed2df49. Before the final lint correction, the resulting tree differs from your reviewed head only by the Wasmtime dependency fix and its changelog. Final commit 896c6e4 removes the non-Unix no-op Result helper and cfg-gates the Unix-only stored directory; Unix sync and Windows write-through rename semantics are unchanged. Feel free to adopt that branch or cherry-pick just the correction after your own rebase. I have not rewritten your fork branch or opened a duplicate PR. Local checks: audit 99 passed / 2 ignored; config 138 passed; kernel retention 2 passed; formatting passed. Kernel Clippy is still running. This is not Windows CI proof. #2008 separately carries the mount cleanup regression/fix. |
91f2742 to
896c6e4
Compare
|
Applied the tested integration to this existing branch with a lease against 91f2742; no intervening author changes were overwritten. Current head is 896c6e4. The description is updated and new fork CI runs are approved. Retention behavior is unchanged; this brings in main and fixes the Windows-only no-op helper lint. The separate mount cleanup fix remains #2008. |
## Linked Issue Closes #2006 ## Summary Keep the private shared mount container alive when retiring an individual lease. Concurrent lease creation can already hold a directory handle to that container; unlinking it makes the next mkdirat fail with ENOENT. This is a pre-existing lifecycle bug investigated after the Ubuntu failure on #2002, not an audit-retention change. The deterministic regression reproduces the same error class; the original CI interleaving was not captured. ## Changes - Remove only the retiring lease endpoint, manifest, and UUID directory, not its shared parent. - Keep the existing private-directory and permission validation. - Add an isolated regression using the real cleanup function and an open parent handle. No live daemon or mount is used. ## Verification - Before the fix: new regression fails with ENOENT at mkdirat. - After the fix: all 17 focused storage-mount tests pass, including maximum-file-I/O framing. - Full astrid-kernel library suite: 598 passed, none failed or ignored. - cargo clippy --locked -p astrid-kernel --lib --tests -- -D warnings: passed. - cargo fmt --all -- --check and git diff --check: passed. - Local execution was on macOS. Linux and Windows validation awaits CI; no native mount certification claimed. ## Test Plan Run the regression and platform CI. Confirm cleanup still removes lease-owned resources and preserves the parent needed by another creator. The retained empty parent is intentional, not a leaked lease. ## AI / Tool Assistance Assisted-by: Codex. Traced the cleanup/creation race, wrote the focused fix and regression, and executed before/after validation. ## Checklist - [x] Linked to an issue - [x] Changelog fragment added - [x] I understand every change and its risks and validation. - [x] I reviewed and tested the tool-generated changes. - [x] Commits have matching Signed-off-by trailers. Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com>
joshuajbouw
left a comment
There was a problem hiding this comment.
Rechecked the integration at 896c6e4. The retention-only commits were applied onto main without changing their behavior; the additional delta is the Wasmtime security update and the Windows-only directory-sync lint correction. I applied that integration, so this is not an additional independent review of my correction.
The original retention review still applies. Local audit/config/retention checks passed (239 tests, 2 ignored), as did kernel Clippy and formatting. Fresh Windows CI now passes, confirming the original lint failure is resolved. The unrelated mount cleanup race is fixed by merged #2008. The remaining MUSL smoke job must finish successfully before landing.
## Linked Issue Closes #1999 Part of #1997. #1996, #2002, and the mount cleanup fix #2008 are merged. Current head `c92562e65087fe47e313e4c029781e89a69927d4` is based on main `26a500a6`. Remaining merge order: this PR, #2004 (recording coverage), #2005 (entry format v2). The integration preserves Pavel's ordered host-audit implementation. The overlapping configuration is resolved in the existing `astrid-config::audit` module, retaining both retention validation and fail-closed host-call settings. Kernel startup loads both settings from the same audit configuration. This PR replaces the host-audit writer. #2004 changes the old writer's fold key to include the capsule. Whichever of the two lands second is rebased onto the other, and the capsule then becomes part of the run key. ## Summary The host-audit sink could lose calls and reorder them without a trace: - It coalesced allowed calls per class over `host_coalesce_ms`, and appended `repeats=N` to the outcome details. The entry signature covers the outcome only as a success bit, so that count was not signed. - It drained pending work per key, so a chain's entries did not follow call order. - A call that met the full queue (`host_queue_capacity`, default 4,096 keys) was dropped, and only a counter moved. - A failed batch was dropped, and calls still queued when the daemon died vanished. As a result, even an externally anchored log could not show that it was complete. This PR records host calls in call order, and makes every loss visible as a signed entry on the chain. It also lets an operator make chosen host-call classes fail closed. ## Changes - **`astrid-audit`: signed host-call records.** These actions are covered by the entry signature: - `HostCallRun`: consecutive calls of one principal. It holds the call count, the first and last call times, a tally per class and outcome, and a fold over every call. - `HostCallLoss`: calls that were accepted but not recorded individually. - `HostCallGap`: an earlier run of the lane stopped without draining. - `HostCallAdmitted`: the write-ahead entry of a fail-closed call. `astrid_audit::host_call` specifies the per-call digest and the fold, so a verifier can recompute them without the kernel. Known-answer tests pin them. - **`astrid-kernel`: ordered lane** (`audit_sink/lane.rs`, `writer.rs`). - Each principal chain has a FIFO of pending slots, so FIFO order is call order. - The writer takes slots from the front and appends each batch atomically with `append_batch_with_principal`. - A failed batch is retried as is, with backoff up to 5 s, before anything behind it is taken. - Signing happens at the durable commit, because only the commit fixes a chain position. The kernel router appends to the same chains directly. - **`astrid-kernel`: lossless coalescing.** - Consecutive allowed and failed calls share one run, and consecutive identical denials share one run. - A one-call run is written as the plain action, as before. - **`astrid-kernel`: loss and gap entries.** - Queue capacity bounds slots, not calls. Overflow folds into one loss slot per chain, which is written as `HostCallLoss` at its place in the chain. - A lane marker in `system:control:audit-lane` records the run and the chains it wrote. After an unclean stop, the next start writes `HostCallGap` into each of those chains and into the session's system chain. - `astrid-storage`: the state-owner resolver admits `system:control:audit-lane` as a system control projection, like `audit`, `invites` and `pair-tokens`. - An unreadable marker, or a marker store that fails, also produces a gap entry. - If the writer thread dies, the lane closes and health reports it. - **`audit.host_fail_closed`** (`astrid-config`, `astrid-capsule`, `astrid-kernel`). - It lists host-call classes whose effect runs only after a write-ahead entry is durable: `file_read`, `file_write`, `file_delete`, `net_connect`, `net_bind`, `process_spawn`. The default is empty. - `HostAuditSink::admit()` is called by the fs, net and process host functions after the security gate and before the effect. - A refusal fails the call with `unknown("audit unavailable")` and is recorded as a denial. - A workspace layer can add classes but not remove them. - **Health:** `accepted` and `persisted` now count calls, and `lost`, `gaps_recorded` and `dropped_after_shutdown` are new. The admin `AuditHealth` wire type is unchanged. - Changelog fragment. **Consequences.** - A denial between allowed calls closes the run. - Order is guaranteed among a chain's host-call entries. Admin rows are still appended when they run. - Calls reported after shutdown are counted in health, not written. **Overhead** (release build, in-memory log, per producer thread; `audit_sink::tests::bench`, ignored by default): | Workload | Before | After | |---|---|---| | 1 thread, 100k mixed allowed calls | 1.61–1.65 µs/call | 0.84–0.92 µs/call | | 8 threads × 8 principals, 200k calls | 15.1–17.9 µs/call | 8.9–12.8 µs/call | | 4 threads, reads with 2% denials | 6.7–7.8 µs/call | 4.2–4.8 µs/call | | 20k distinct denials, 1 thread | 1.35–1.58 µs/call; 15,904 calls dropped | 1.82–1.91 µs/call; none dropped, overflow in 5–7 loss entries | ## Verification Current integrated tree: full kernel library **628 passed, 1 ignored**; capsule audit tests **45 passed**. Earlier validation of the same recorder/retention integration: audit 106 passed / 2 ignored, config 140 passed, kernel audit filter 51 passed / 1 ignored; audit/config/kernel library Clippy passed with warnings denied. The published tree is byte-identical to the tested combined tree; CI must pass on this new head before merge. Prior author verification (before this integration): - **Test suites:** `astrid-audit` 74, `astrid-capsule` 748, `astrid-config` 139 and `astrid-storage` 864 all pass. The kernel lib passes 610 of 611 with `--test-threads=6` and umask 0022. The remaining failure needs a copy-on-write workspace backend and also fails on `main`. - **Unit and behaviour tests:** - host-call digest and fold known-answer values; every field changes the digest; - a dropped, reordered or altered call fails `matches_calls`; - changing a signed count breaks the signature; - call order across 6 producers, 3 principals, 8-entry batches and a 10 ms window (the same probe fails on `main`); - a 40-call run whose fold, count and time range recompute from the calls; - queue overflow written as one loss entry that commits to the lost calls; - gap entries after an abandoned lane run, none after a clean restart, and gap duties carried across a second unclean stop; - a corrupt marker and a failing marker store, including the unavailable-marker gap surviving a later marker write; - a system-only gap recorded at start with no host call, and retried while idle after a refused append; - the marker namespace written and read back through a runtime principal store; - a batch the log refuses, retried and written in order once the cap is raised; - a fail-closed write durable before `admit()` returns; - refusal within the first failed batch, and after shutdown; - config validation and workspace restriction; - `connect-tcp` and `bind-tcp` refusing without a connection attempt. - **Lint:** `cargo fmt` and `cargo clippy --all-features --all-targets -- -D warnings` on the touched crates. ## AI / Tool Assistance Assisted-by: Claude Code:claude-opus-5-5. Covers design, implementation and tests in all touched crates, and this description. Assisted-by: Codex CLI. Review passes over the diff. Current integration and verification assisted by Codex. Fresh CI pending; no claim of merge readiness until those checks complete. ## Checklist - [x] Linked to an issue - [x] Changelog fragment added under `changes/{issue}.{kind}.md` (docs/CI-only may skip; release PRs roll fragments into the version section instead of adding one) - [x] I understand every change in this PR and can explain its design, risks, and validation. - [x] I reviewed and tested any meaningful tool-generated output included in this PR. - [x] Every non-bot, non-merge commit has a matching `Signed-off-by` trailer. Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com> Co-authored-by: Joshua J. Bouw <jjb@unicity-labs.com>
… changes and capsule loads (#2004) ## Linked Issue Closes #1998 Part of #1997. #1996, #2002, #2003, and #2008 are merged. Current head `c0f4052305da424d9625c9f96332448667a9259b` integrates this PR with main `e63f65cc`. Remaining order: this PR, then #2005 (entry format v2). The integration retains #2003's ordered, loss-accounted writer and fail-closed admission. Capsule identity is part of the run key, so alternating capsules cannot fold together. HTTP/tool/approval coverage records remain individual queue entries; overflow still records a payload-bound loss summary. Existing unattributed call digests keep their exact v1 encoding; attributed and new coverage calls use the documented `astrid.audit.host-call.v2` digest domain. This is a call-fold encoding distinction, not #2005's entry-format change. Ordering boundary: queued records retain per-principal FIFO order. Synchronous HTTP/approval commits still use the existing durable append path, like direct administrative audit writes; no new global ordering guarantee between that path and queued records is claimed. ## Summary The audit log did not record the events that determine what an agent did: - LLM and other HTTP exchanges appeared only in `tracing::debug!`, so an agent turn, including its model request, added no entry; - capsule tool calls and approval decisions had no producer; - capability grants and revocations were recorded as `AdminRequest` rows at authorization time, before the change was applied and without its result, and grant-on-use was not recorded at all; - capsule install and load were not recorded, and install hooks ran without an audit sink; - host-call entries did not name the capsule, and `FileWrite` stored a zero hash. This PR records all of these on the signed log, without changing the v1 entry format. Content never enters the log, only commitments to it. Provider credentials can now be injected by the host, so they never enter guest memory. ## Changes - **`astrid-audit`: entry kinds.** - `HttpRequest` (pre-commit) and `HttpResponse` (completion). - `CapabilityChanged`, `CapsuleInstalled` and `CapsuleLoaded`. - A `CapsuleActor` (capsule id and wasm hash) on file, network, process, tool-call and approval entries. - Linking fields on `CapsuleToolCall` and the approval entries. New fields are optional and omitted when unset, and new variants are appended after the existing ones. Entries written before this change re-serialize to the same signed bytes. - **HTTP exchanges** (`astrid-capsule` HTTP host, `astrid-kernel`). - Each wire request, each redirect hop included, gets a durable `HttpRequest` pre-commit before it is sent. The pre-commit holds the method, host, port, BLAKE3 commitments to the path, headers and body, the redirect hop, and a per-principal sequence number scoped to the kernel run. - An `HttpResponse` completion links to its pre-commit. It carries the status, provider request ids, and a hash of the response body as it was read (incrementally for streams). - Completions are appended durably. - Refusals by the scheme check, egress, the security gate or the SSRF airlock are recorded as denied requests and take a sequence number. - Credential headers, credential query parameters and every secret value the capsule received are redacted before hashing. The redaction rules are documented so a verifier can recompute the commitments. - Recording is best-effort, like the rest of the audit log since 0.9.0. If the pre-commit cannot be written, the request is still sent, a security event is logged, and the missing entry shows as a gap in the run's sequence numbers. Approval prompts behave the same way. Making these fail closed is an operator policy; it would fit as an `http` class of `audit.host_fail_closed` from #2003 once both have landed. - **Host-side credentials.** - A header value can name a manifest-declared secret as `{{secret:NAME}}`, and the host substitutes it on the wire. - Undeclared names are refused and recorded. Placeholders are stripped on a cross-origin redirect. - Commitments cover the placeholder form, and the entry lists injected secret names, never values. - `docs/models.md` describes how a provider capsule adopts this. - **Tool calls and approvals.** - One `CapsuleToolCall` entry per `tool.v1.execute.<tool>` invocation, with the call id, argument and result hashes, and the capsule. - Approval prompts are committed before they are published. Every decision is linked to its prompt, with how it was reached: user, session allowance, remembered consent, timeout and so on. - Grant-on-use prompts are committed by the dispatcher before `GrantRequired` is published. An approval is committed before `GrantResult` is published, so the chain shows it before the grant it caused. - **Capability changes.** Applied changes are recorded on the affected principal's chain after they are saved: token mint and revoke, caps grant and revoke (only the patterns actually added), group and capsule membership changes, distro self-grants, and grant-on-use. - **Code identity.** - `CapsuleInstalled` binds the wasm and exact `Capsule.toml` hashes. - `CapsuleLoaded` (on load or replace) binds the verified wasm hash, the manifest hash and the engine profile. - Install and upgrade hooks run by the daemon get the signed sink, attributed to the hook's component. - **Attribution.** - Host-call entries name the capsule and wasm hash, stamped by the host from the verified component; a guest cannot choose them. - `write-file` reports the hash of the bytes written. - Changelog fragment. ## Verification Current signed integration head: **111 audit tests passed / 2 ignored; 777 capsule tests passed; 643 kernel tests passed / 1 ignored** (1,531 passed in total). Commands: `cargo test --locked -p astrid-audit -p astrid-capsule -p astrid-kernel --lib -- --quiet`; `cargo check --locked --workspace`; `cargo clippy --locked -p astrid-audit -p astrid-capsule -p astrid-kernel --lib --tests -- -D warnings`; formatting and diff checks. All passed locally on macOS. Added integration regressions bind capsule/wasm identity in folded digests and prove coverage events remain individual while overflow binds the original payload. The attribution test now asserts FIFO order across alternating capsules instead of expecting the old unordered folding behavior. Fresh platform CI is still required. Prior author verification before integration: - **Test suites:** `astrid-audit` 71, `astrid-capsule` 777, `astrid-capsule-install` 106 and CLI 818 all pass. The kernel lib passes 601 of 602 with `--test-threads=6` and umask 0022. The remaining failure needs a copy-on-write workspace backend and also fails on `main`. - **Unit and integration tests:** - legacy JSON encodings pinned, and a legacy entry re-verified after a round trip; - pre-commits resolve before a loopback server sees the request; - buffered and streamed responses hashed and linked, including streams closed early, transport failures and gate denials; - redaction: short secrets, many secrets, overlapping secrets; - credential injection end to end: the server receives the secret, the guest never reads it, and the commitment covers the placeholder; - placeholder rules: undeclared, malformed and unterminated placeholders refused, also after a blank secret, and an omitted header names no secret; - tool results captured through the IPC host function, with error, missing, mismatched and cancelled cases; - approval prompts committed before they are visible on the bus; - capability grant, revoke and membership entries, with none for a no-op re-grant, and a grant-on-use approval ahead of its grant; - a daemon install producing `CapsuleInstalled` and `CapsuleLoaded` with the expected hashes, and an upgrade recorded as load then replace. - **Lint:** `cargo fmt` and `cargo clippy --all-features --all-targets -- -D warnings` on the touched crates. `cargo check --target wasm32-unknown-unknown` of `astrid-audit`, `astrid-capsule` and `astrid-kernel`. ## AI / Tool Assistance Assisted-by: Claude Code:claude-opus-5-5. Covers design, implementation and tests in all touched crates, and this description. Assisted-by: Codex CLI. Review passes over the diff. Integration onto the landed recorder and associated regression tests assisted by Codex. The earlier source review does not constitute independent review of these integration edits. ## Checklist - [x] Linked to an issue - [x] Changelog fragment added under `changes/{issue}.{kind}.md` (docs/CI-only may skip; release PRs roll fragments into the version section instead of adding one) - [x] I understand every change in this PR and can explain its design, risks, and validation. - [x] I reviewed and tested any meaningful tool-generated output included in this PR. - [x] Every non-bot, non-merge commit has a matching `Signed-off-by` trailer. --------- Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com> Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com> Co-authored-by: Joshua J. Bouw <jjb@unicity-labs.com>
…d-key verification (#2005) ## Linked Issue Closes #2000 Part of #1997. The current integration includes the landed #1996, #2002, #2003 and #2004. Proposed merge order: #1996, #2002, #2003 (ordered recording), #2004 (recording coverage), this PR. v2 derives an entry's sections from the serde form of the action, authorization and outcome. The new action variants in #2003 and #2004 therefore encode without any change to this format or its specification. ## Summary Format v1 cannot be verified outside Astrid with confidence: - its signed bytes mix binary fields with `serde_json` of the action and authorization, and drop either one if serialization fails; - it signs whole-second time, and only a success bit of the outcome; - it has no sequence number; - verification trusts the public key embedded in each entry, so anyone with write access to the store can re-sign a rewritten chain under any key and still pass; - one runtime key signs audit entries, capability tokens and builds. This PR adds format v2, which is off by default and enabled with `[audit] entry_format = "v2"`: - a deterministic CBOR body that signs every stored field; - a per-chain sequence number; - a dedicated audit key; - verification against a cross-signed key registry instead of the embedded key. v1 history is kept as it is, and the v2 chain links to it. ## Changes - **`astrid-crypto`:** `verify_strict` (RFC 8032 strict Ed25519) and a `random_bytes` helper. - **`astrid-audit`: `entry_v2`.** The module documentation is the byte-level specification, with a known-answer test that an implementation written from the specification alone reproduces. - The body (RFC 8949 §4.2.1 deterministic CBOR) carries: - the format tag; - a chain id derived from the registry, session, principal UID and alias; - a sequence number and the previous hash; - nanosecond time, entry and session ids, and the acting capsule; - the action, authorization and outcome as `[tag, {field => commitment}]` sections derived from their serde form; - the signer key and its key epoch. - Field values are salted commitments. Salts are HMAC-SHA256 of a per-entry salt key, so one field can be disclosed without the others. - The entry hash is SHA-256 of the body, and the audit key signs a domain-separated wrapper of it. - **`astrid-audit`: key registry.** - A hash chain of records binding keys to roles: audit, capability, build and audit-v1. - The genesis is signed by every key it binds, and a rotation by both the old and the new key. A key can never be registered twice. - **`astrid-audit`: verification.** `ChainVerifier` accepts a v2 entry only if all of these hold: - its signer holds the audit role at the entry's epoch; - its chain id derives from the registry; - sequence numbers run 1, 2, 3; - links hold; - epochs never decrease. With a registry present, v1 entries and archive receipts must be signed by registered keys. v2 receipts carry the signer's epoch. - **`astrid-audit`: writing.** - `AuditLog::enable_entry_v2` writes the genesis on first use. - The first v2 entry of each chain has sequence 1 and links to the chain's last v1 entry. - After that, v1 appends are refused (`V1Closed`). - Storage enforces the v1 closure and the current key epoch at commit time, under the durable append lock, so another opener's switch or rotation cannot be bypassed. - The same check applies to a prune before anything is deleted: a new deletion plan is accepted only if its receipt's signer is the registry's current audit key (a receipt without an epoch only while no registry exists). A plan accepted earlier still finishes after a rotation. - `rotate_audit_key` appends a cross-signed rotation. - **`astrid-config`, `astrid-kernel`: `[audit] entry_format`.** It is operator-only, and the default is `"v1"`. - With `"v2"`, the kernel loads or creates `keys/audit.key` (owner-only), enables v2, and on first use writes a genesis that registers the runtime key for the capability, build and v1-audit roles. - A store with a registry stays on v2. - Boot stops if the registry does not verify, if the audit key is missing or different, or if the configuration cannot be read on a store that is not yet on v2. - `keys/audit.key` is protected from admin filesystem edits. - It applies on every native host. - **Docs and changelog.** - `docs/config.md` documents the switch and its one-way migration. - Rolling back to a release without v2 is not supported. - Changelog fragment. With v2 off, v1 behaviour, verification results and archive keys are unchanged. ## Verification ### Current stack integration Head `0dd52fdd48fd25eab5c778353ecadd9701033ff8` reconciles landed #2004 (`3b3360fd`) with integration `c225dfc7`. Its tree is byte-identical to `c225dfc7` (tree `c81645720ec1d81d2c1e447aecdf328f54b4e0a0`). All Actions jobs, including both MUSL smokes, passed on that tree. Three CodeQL hard-coded-value alerts were individually investigated and dismissed as false positives: an HMAC output buffer fully overwritten before return, a test-only known-answer key, and an OS-randomness buffer that cannot return on RNG failure. No scanner rule or test was weakened. Required checks on the ancestry-only commit must still complete before merge. - Retains both anchor-watermark and signing-key-epoch checks before prune-plan acceptance, as well as ordered host recording and capsule attribution. - Keeps `entry_format`, retention settings and fail-closed settings in the shared `AuditConfig`, with both operator-only restrictions intact. - Adds regressions for v2 pruning at the anchor boundary and commitment to host-stamped capsule identity. - The forged-receipt test still requires signature rejection. It now restores the authentic receipt before its subsequent legitimate-prune scenario, because immutable receipt history correctly rejects replacing that generation. - Audit: 171 passed, 2 ignored. Config: 143 passed. Crypto: 28 passed. Kernel: 648 passed, 1 ignored with `--test-threads=4`. The first parallel kernel run hit two unchanged 2-second response timeouts; all three tests in that group passed unchanged in an isolated rerun (0.71s), then the complete four-thread kernel run passed (51.72s). - Audit doctests: 2 passed. Workspace `cargo check --locked --workspace`, formatting and focused audit/config/crypto/kernel Clippy (`--lib --tests -- -D warnings`) passed. - `cargo check --locked -p astrid-kernel --target wasm32-unknown-unknown` passed, with unused/dead-code warnings; this is build evidence, not a browser runtime test. - The prior independent review applies to the earlier source head. These integration checks do not claim a new independent review; completed CI on the identical earlier tree is distinguished from checks on the current commit. ### Original author validation (before integration) - **Test suites:** `astrid-audit` 123 (+2 doc), `astrid-config` 140 and `astrid-crypto` 31 all pass. The kernel lib passes 592 of 593 with `--test-threads=6` and umask 0022. The remaining failure needs a copy-on-write workspace backend and also fails on `main`. - **Unit and integration tests:** - the known-answer test, and CBOR against RFC 8949 vectors, with non-canonical input rejected; - every stored field covered by the signature: each JSON leaf of a stored entry, and each field of every action, authorization and outcome variant; - sequence gaps, including an entry deleted from the store; - unregistered, foreign, retired and backdated keys rejected, v1 after v2 rejected, and epoch regression across a chain change detected; - registry genesis, rotation and tamper rules; - v1 to v2 migration with mixed chains, persistence and rotation across restart, and principal UID binding; - concurrent single and batch appends producing gap-free sequences; - commit-time refusal of v1 after another opener enables v2, and of a stale epoch after another opener rotates; - archive receipts bound to a registry epoch, and forged receipts rejected; - a stale handle's prune refused, with entries, receipts and verification unchanged: one without v2 after another enabled it, and one at epoch 0 after another rotated to epoch 1; a plan accepted before a rotation finishes after it; - kernel boot: v1 default, key creation, v2 persistence with the setting back at v1, and a missing or replaced key or unreadable config stopping boot. - **Scratch daemon:** boot on v1, switch to v2, then restart with the setting back at v1. The chain reads 8 v1 entries followed by v2 entries 1 to 10, bound to the principal UID, under one registry. It verifies with registered v1 keys required. - **Lint:** `cargo fmt` and `cargo clippy --all-features --all-targets -- -D warnings` on the touched crates. `cargo check --target wasm32-unknown-unknown` of `astrid-kernel`. ## AI / Tool Assistance Assisted-by: Claude Code:claude-opus-5-5. Covers design, implementation and tests in all touched crates, the format specification, and this description. Assisted-by: Codex CLI. Review passes over the diff. Merge after #2004 and successful checks on the current integration. ## Checklist - [x] Linked to an issue - [x] Changelog fragment added under `changes/{issue}.{kind}.md` (docs/CI-only may skip; release PRs roll fragments into the version section instead of adding one) - [x] I understand every change in this PR and can explain its design, risks, and validation. - [x] I reviewed and tested any meaningful tool-generated output included in this PR. - [x] Every non-bot, non-merge commit has a matching `Signed-off-by` trailer. --------- Signed-off-by: Pavel Grigorenko <pavel@unicity-labs.com> Signed-off-by: Joshua J. Bouw <jjb@unicity-labs.com> Co-authored-by: Joshua J. Bouw <jjb@unicity-labs.com>
Linked Issue
Closes #2001
#1996 is merged. The retention-only commits are rebased onto main
ed2df49b, preserving the Wasmtime 48.0.3 security update. Current head:896c6e4c3d9f5fb305278a3a408250a7908d7204.The integration adds one Windows Clippy fix: directory syncing and its stored path are compiled only on Unix; Windows retains the existing write-through rename behavior. No global lint suppression.
Summary
Today, pruning deletes entries permanently. That holds whether an operator runs it or the global cap triggers it (1,000,000 entries / 1 GiB), and only a chain's latest prune receipt is kept. So history no external party has seen can be erased, and a verifier cannot connect retained entries to history it anchored several prunes earlier.
This PR changes that:
Installs that do not anchor keep today's bounded retention.
Changes
astrid-audit: anchored watermarks.AuditLog::mark_anchoredrecords that a chain was certified through a position. The position counts entries from the chain's genesis, pruned ones included (omitted_total + count). The mark carries the head hash at that position and up to 4,096 bytes of opaque evidence.position - 1, or, at the pruned boundary, the latest receipt's terminal hash, must hash to the given head.audit:anchor_marks) and are replaced by compare-and-swap. Chain metadata, which append-intent recovery compares byte for byte, is untouched.astrid-audit: prunes never pass a watermark.prune_oldest(used byaudit pruneand the cap) takes the oldest sealed segment, in seal order, whose removal is allowed. It searches the whole seal-order index before refusing.astrid-audit: the cap holds instead of deleting. When an append reaches the cap and every candidate holds unanchored history:The hold clears when a watermark advances, a prune finishes, a segment is sealed or the caps change.
astrid-audit: every receipt is kept. Receipts are recorded underaudit:receipt_history/<chain>/<generation>before installation. A chain pruned before this change keeps its latest receipt.AuditLog::prune_receiptspages them.astrid-audit: archive, then drop. With anAuditArchiverset, a prune writes the entries it removes to an archive and commits it before the deletion plan is created. The archive must match the receipt's count, digest and terminal hash, or the prune is abandoned. Prunes are serialized from segment selection to plan installation, so two prunes of one chain cannot archive the same generation.astrid-core,astrid-kernel: admin methods.audit.anchor_mark(capabilityaudit:anchor) accepts one certification's evidence and up to 4,096 chains. Each chain is reported as advanced, unchanged or rejected with the reason. Malformed evidence, an empty or oversized list, or a duplicate chain rejects the whole request.audit.headsandaudit.exportdo. Rejections and denials are recorded.audit.anchor_status(capabilityaudit:heads) reports each chain's count, pruned total, watermark, its time and evidence, whether anchoring is required before prune, and the retention hold.audit.exportgainsreceipts_from, which pages prune receipts with their hashes and signing bytes.audit.statsreports the hold.astrid-config:[audit.retention], operator-only.require_anchor(defaultfalse) makes every prune require a watermark.archive_dirmust be an absolute path. It writes each pruned segment as<session>/<system|principal.alias>/<generation>.jsonl: the signed receipt, then the removed entries. The file is written under a temporary name, synced and renamed, in directories private to the daemon's user.AuditConfigmoves to its own module.astrid-cli:astrid audit anchor-statusshows per-chain watermarks and lag, and exits 2 while the hold is set.audit pruneexplains a refusal.audit statsshows the hold.audit export --receipts-fromprints prune receipts.Changelog fragment.
New storage fields are optional and absent by default, so existing metadata re-encodes to its stored bytes.
Verification
Current integration validation: audit library 99 passed / 2 ignored; config library 138 passed; kernel retention 2 passed; kernel-library Clippy with warnings denied, formatting and diff checks passed on macOS. All 33 checks on 896c6e4 now pass, including Windows and both MUSL targets. #2008 is merged. The exact combined tree with main bdc1fe8 (tree 6c15a1dd94a8247ff6181a27cdb15a4cb5f77a98) also passes the full kernel library suite: 605 passed, zero failed or ignored.
Prior author validation of the retention implementation follows:
astrid-audit99 (+2 doc),astrid-core390,astrid-config138,astrid-uplink58 and CLI 818 all pass. The kernel lib passes 602 of 603 with--test-threads=6and umask 0022. The remaining failure needs a copy-on-write workspace backend and also fails on the base commit. With full parallelism on a loaded host, some kernel-router tests exceed their 2-second admin-response timeout; feat(audit): signed audit.heads snapshot and paged audit.export #1996's branch shows the same under the same load.anchor_markoutcomes and whole-request rejections;cargo fmtandcargo clippy --all-features --all-targets -- -D warningson the touched crates.AI / Tool Assistance
Assisted-by: Claude Code:claude-opus-5-5. Covers design, implementation and tests in all touched crates, and this description.
Assisted-by: Codex CLI. Review passes over the diff.
#1996 and #2008 are merged; all current-head CI checks pass.
Checklist
changes/{issue}.{kind}.md(docs/CI-only may skip; release PRs roll fragments into the version section instead of adding one)Signed-off-bytrailer.