Skip to content

Wave 2a: TXN key isolation, AOF replay clock + everysec fsync agent, graph WAL overflow, expiry/ACL parity, held-file release - #1316

Merged
TinDang97 merged 99 commits into
mainfrom
claude/gifted-mendel-e9wiz5
Oct 4, 2026
Merged

TinDang97 merged 99 commits into
mainfrom
claude/gifted-mendel-e9wiz5

Conversation

@TinDang97

Copy link
Copy Markdown
Collaborator

Summary

Wave 2a of the v0.9.2 perf review (plan: .add/milestones/v0-9-2-perf-review/plans/WAVE2-PLAN.md, stage 1 plus three adversarial review rounds). It fixes eight issues and closes the gaps the reviews found in the fixes themselves:

  • moon#1299 / moon#1303 — TXN isolation (KV plane). A key written inside an open cross-store TXN is held until COMMIT/ABORT.
    • Other clients' writes to it, and FLUSHDB/FLUSHALL/SWAPDB of its database, answer -TXNCONFLICT.
    • Eviction and active expiry skip held keys.
    • Every connection exit rolls the TXN back before the socket closes. That covers protocol faults, blocked pops whose peer vanished, output-limit closes, subscriber QUIT/fault, PSYNC and RESET; the cleanup runs in one wrapper around each runtime's connection handler.
    • New INFO gauges: txn_open, txn_oldest_age_ms, txn_held_keys, txn_conflicts_refused.
    • A TXN.COMMIT refused after KILL SNAPSHOT now rolls its writes back.
  • moon#1302 — graph/MQ/workspace/temporal WAL records past the 4096-slot append channel. These were acked and then lost on kill -9. They now wait in a shard-owned overflow queue that the 1 ms tick drains in order. A large TXN ABORT graph rollback is written to the WAL before it is replicated, and only the accepted prefix is replicated.
  • moon#1283 — AOF expiry replay clock. The writer emits MOON.TS <ms> stamps and replay pins the clock to them; it no longer trusts the log's mtime.
    • A clean stop ends each incr with MOON.TS <ms> CLOSE. Records an older binary appends after a downgrade are recognised by their position and judged consistently on every later boot. There is no mtime heuristic and no forced rewrite.
    • A stop that could not write its marker logs that it could not, instead of "drained and synced".
    • The procedure for the unprotected case (a downgrade after an unclean stop) is in STORAGE-FORMAT §3.3 and the production guide.
  • moon#1266 (Option 3) — appendfsync everysec vs kill -9.
    • The fsync runs on a per-writer agent thread. Its one-word hand-off has a loom model, including the agent's unwind path.
    • The tokio writer flushes every batch; the monoio writer warm-polls while writes flow and parks when idle.
    • Lossy reps under kill -9: 226/240 → 9/240 (16/360 across both tokio runs). The remaining window is Option 1A (WS46).
    • A stalled fsync is now reported the redis way: the "Asynchronous AOF fsync is taking too long" log line, INFO aof_pending_bio_fsync and aof_fsync_in_flight_ms, and aof_delayed_fsync at redis's 2 s cadence.
  • moon#1286 — INFO expired_keys now counts what redis 7.2.7 counts, compared case by case against the oracle:
    • active and lazy expiry;
    • writes over an expired key, cold-only keys included;
    • no recount during AOF replay;
    • absolute deadlines already in the past delete the key and publish del, and the master propagates that delete as DEL;
    • CONFIG RESETSTAT.
  • moon#1296 — ACL command rules render in application order (redis 7.2+). ACL SAVE / ACL LOAD round-trips are now exact.
  • moon#1289 — held cold spill files are released automatically.
    • With an AOF, through a fold. Without one, through a rate-limited snapshot, which also runs with save "".
    • The snapshot never contains an open TXN's uncommitted writes. It waits while a TXN is open, and it is abandoned and retried if a shard holds a TXN write at its start.
  • Bugs that were already on main, found and fixed during review:
    • after a restart, AOF writes to db 0 replayed into the previous run's last db;
    • CONFIG SET appendfsync was not applied to the writers;
    • the everysec fsync-error status cleared itself;
    • volatile-* eviction answered OOM while a TTL key existed (sampling miss); it now falls back to the nearest deadline.

README: the everysec "kill-9-lossless" claim now states the measured window, and the v0.8.0 row says "zero synced-write loss".

Checklist

All results below are from a Linux container (4 vCPU), not the merge bar. scripts/ci-local.sh cannot run here.

  • cargo fmt --check passes.
  • cargo clippy --all-targets -- -D warnings passes (monoio), and the same with --no-default-features --features runtime-tokio,jemalloc.
  • cargo check --manifest-path fuzz/Cargo.toml --all-targets. New fuzz target aof_incr_replay is in both fuzz.yml matrices.
  • cargo test --all-features: not run (gpu-cuda cannot build here). Instead:
    • cargo test --release --lib: monoio 7057 passed, 1 failed. The failure is the pre-existing statistical LFU test lfu_reads_grow_the_frequency_and_an_overwrite_keeps_it (0.15% tail; 0/60 on rerun). Tokio 6111 passed, 0 failed.
    • Integration: 74 suites × 2 runtimes, --include-ignored, MOON_BIN pinned to release builds of bda76c1. All green except the known reds listed under Notes. The two later product commits were re-gated on their own: the crash_recovery_cold_del_rewrite test rework at ≥8/8 on both runtimes, and the ColdIndex::remove early return with storage lib tests plus 13 cold-tier/TXN/expiry suites on both runtimes.
    • Loom loom_aof_fsync_agent: 5/5 (cargo rustc --release --test loom_aof_fsync_agent -- --cfg loom).
  • Consistency tests (./scripts/test-consistency.sh): not run in this container. The parity-sensitive changes were byte-compared against a redis 7.2.7 oracle instead (expired_keys, keyspace events, GETEX errors, ACL render/round trip).
  • Hosted matrix (Windows, MSRV 1.94, memory steady-state): to be dispatched on this PR.
  • Review: three adversarial review rounds over the integrated tree, one reviewer per area (TXN/WAL, AOF, parity/held files). The final pass reports nothing above NIT in this PR's code.

Performance Impact

Hot paths are touched: the write dispatch (the moon#1299 hold check, the moon#1286 overwrite accounting) and the AOF writer (moon#1266, moon#1283).

Setup. A/B of main cf6fa65 vs wave 2a bda76c1, both cargo build --release. 4-vCPU Linux container, quiet box, client on the same host. The A and B runs are interleaved, each on a fresh server. Values are the median with (min–max) of the reps.

Throughput (redis-benchmark, rps). 3 reps per cell, 7 reps for the cells marked †.

cell monoio main → 2a tokio main → 2a
s1, no AOF, GET p1 c50 104.5K (103.5–109.3) → 106.6K (98.4–118.6) +1.9% 111.2K → 112.9K +1.6%
s1, no AOF, SET p1 c50 104.0K → 110.9K +6.6% 112.0K → 114.4K +2.1%
s1, no AOF, SET p16 c50 † 978K (789–1044) → 939K (905–1015) −4.0%, inside base range 408K → 398K −2.6%
s1, everysec, p1 c1 20.5K → 20.4K −0.4% 16.2K → 16.2K 0.0%
s1, everysec, p1 c50 103.4K → 113.0K +9.3% 41.1K (38.6–42.6) → 54.1K (49.1–54.3) +31.6%
s1, everysec, p16 c1 235.6K → 235.8K +0.1% 114.0K (107.5–126.6) → 161.4K (156.6–161.7) +41.6%
s1, everysec, p16 c50 784K → 797K +1.6% 163K (159–219) → 257K (256–263) +57.9%
s4, everysec, p1 c50 † 61.8K → 63.4K +2.4% 61.4K (59.5–68.8) → 59.4K (56.8–63.0) −3.3%, ranges overlap
s4, everysec, p16 c50 654K → 680K +4.1% 498K → 473K −5.0%, ranges overlap
s1, always, p16 c50 † 472K → 477K +1.1% 402K → 417K +3.5%

The tokio everysec win at s1 comes from moon#1266: the writer flushes every batch and the fsync runs on its own agent thread. The other deltas fall inside the run-to-run spread. On this box, two identical binaries differ by about ±6% per comparison.

Write-path cost per op.

  • An earlier 5-rep server-CPU sample suggested +10.6% CPU/op for no-AOF SET p16. A profile pass did not confirm it:
    • Callgrind (shard thread, 300k SET p16, --shards 1, offload on): main 2795.9 → wave 2a 2806.9 Ir/SET. That is +11 instructions, +0.4% (the moon#1286 compare plus the moon#1299 any_held() gate). Cache misses and branch mispredicts are unchanged.
    • Native CPU/op, 12 rotating reps: main 985 ns → wave 2a 980 ns (−0.5%, 95% CI −5.0%…+7.1%).
    • Tokio, 8 reps: main 2435 ns → wave 2a 2355 ns.
  • The same profile found an older cost and fixed it (ColdIndex::remove early return): every overwrite hashed the key into an empty cold index. With the fix the write path runs 2605.9 Ir/SET, −7.2% vs both main and wave 2a. Native: 960 ns, −2.5%.

Durability (moon#1266, kill -9 1 ms after the last ack, 20 reps per cell): lossy reps 226/240 → 9/240, or 16/360 across both tokio runs.

Not measured:

  • Hosted or GCE numbers. All figures above are relative evidence from one container.
  • scripts/bench-compare.sh against redis was not re-run.

Notes

Known reds, all also red before this PR:

  • Tokio has no master-side PSYNC: replica_past_deadline_1286, replication_ttl_semantics, and the replica tests in scripts_in_multi_894 / script_move_copy_db_1068.
  • Tokio builds have no graph engine: crash_recovery_graph_durability.
  • aof_fold_exactly_once_455::exec_parked_in_wait_across_a_rewrite_replays_once_toplevel.
  • Tokio cold_tier_aof_double_apply_902.
  • The monoio parked_idle_parity::unauthenticated_conn_never_task_parks timing gauge.
  • Two tests in review_w1_txn_abort_no_aof_snapshot_1285 stay red until WS42 (moon#1300, wave 2b).

Design decisions taken by the maintainer:

Trade-offs:

  • An open, or constantly busy, TXN keeps held cold files on disk and SWAPDB refused. This shows in cold_held_release_snapshots_deferred_txn and ..._abandoned_txn.
  • Every clean AOF stop adds about 55 bytes per incr.

Follow-ups filed: moon#1306–#1315. The most serious are moon#1314 (an AOF write-error latch keeps acking writes under everysec), moon#1309 (no -MISCONF refusal) and moon#1307 (graph-plane TXN isolation).

Also in this PR, not wave-2a code: the ColdIndex::remove early return (perf, above), found while profiling.

Per-workstream evidence is in .add/milestones/v0-9-2-perf-review/plans/{WS36..WS41,R1-fix-*,R2-fix-*,R3-fix-*}/SUMMARY.md.

🤖 Generated with Claude Code

https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu


Generated by Claude Code

Summarises round 1 (#1292), wave 1 (#1301) and its follow-up (#1305): what
shipped per issue, measured costs, the five review rounds and why they point
at #1299, and the open residuals. Plans wave 2 as a DAG of 11 workstreams in
three lanes (TXN, AOF stream, cold/expiry/parity) over two stages with an
adversarial review + integration gate after each, and lists the maintainer
decisions needed before starting.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… oracle)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… TXN (moon#1285)

A BGSAVE taken while a TXN is open holds its uncommitted writes (by
design: point-in-time, needed by the AOF fold). With --appendonly no the
abort's compensating records have no log to go to, and the snapshot is
the durability authority: kill -9 + restart brings the aborted writes
back.

Pre-existing: red on f7f1d96 monoio and tokio, and on ce65400 monoio. The
WS27 durability claim does not cover this quadrant.

MOON_BIN=... cargo test --test review_w1_txn_abort_no_aof_snapshot_1285 -- --include-ignored --test-threads 1

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… AOF recovery (moon#1285)

A crash is an implicit abort, but a TXN's writes reach the AOF as they
run, with no transaction markers, and recovery rolls nothing back.
Pre-existing: red on f7f1d96 monoio and tokio, and on ce65400 tokio.

MOON_BIN=... cargo test --test review_w1_txn_abort_no_aof_snapshot_1285 a_crash -- --include-ignored

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…e, now durably (moon#1285)

A non-TXN client's SET to a key the open TXN wrote is accepted (no
write-write conflict), and the abort restores the pre-TXN value over it.
Live this is pre-existing (ce65400 answers orig too); on ce65400 the
restart replayed the other client's SET and brought it back, on f7f1d96
the compensating RESTORE makes the lost update durable and replicated.

Red on f7f1d96 monoio/tokio (live orig, restart orig) and ce65400 monoio
(live orig, restart other-client).

MOON_BIN=... cargo test --test review_w1_txn_abort_no_aof_snapshot_1285 an_abort_does_not -- --include-ignored

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ture back (moon#1303)

Mechanism: the connection leg of both runtimes captured a key's pre-image
and recorded a write intent BEFORE dispatch and kept both even when the
command answered an error and wrote nothing (`SET k v BADOPT`, `INCR` of a
non-number, WRONGTYPE). `TXN ABORT` then restored the stale pre-image over
whatever another client had written meanwhile - logged to the AOF and
replicated.

The capture now lives in one shared helper, `transaction::conn_capture`
(both handlers call it; `handler_monoio/mod.rs` -52 lines,
`handler_sharded/mod.rs` -21). `capture_conn_write` returns a mark (undo
length + each key's previous intent, which `KvWriteIntents::record_write`
already returns); `ConnWriteCapture::finish` runs after dispatch and, on an
error reply, truncates the undo log back to the mark (`UndoLog::truncate`)
and puts every key's previous intent back (`KvWriteIntents::restore`, new),
newest first so a key captured twice by one write ends where it began. The
pattern of the script leg (`txn_undo_capture` / `txn_undo_discard`, PR #1301).

Evidence: unit tests `transaction::conn_capture::tests::*` and
`kv_mvcc::tests::restore_undoes_record_write`. The real-server test
(`an_erroring_write_holds_nothing_{1,4}_shard(s)` in
tests/txn_isolation_1299.rs: A `TXN BEGIN; SET k v BADOPT`, B `SET k fresh`,
A `TXN ABORT`, `GET k` = fresh; also INCR of a non-number and two WRONGTYPE
shapes) lands with moon#1299 in the next commit, which also makes the
capture take back the key hold it adds. Red on base 2e99254+3 (both
runtimes: "TXN ABORT restored the pre-image ... over another client's
write"), green after, monoio and tokio, --shards 1 and 4.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…(moon#1299)

Mechanism: another client's write to a key an open cross-store TXN had
written was acknowledged, and the TXN's ABORT then restored its pre-image
over it (live, after a restart, on replicas). A key the TXN writes is now
HELD until the TXN commits or aborts; any other writer is refused with
`-TXNCONFLICT key held by an open transaction`, and FLUSHDB / FLUSHALL /
SWAPDB touching a database with held keys with `-TXNCONFLICT database has
keys held by an open transaction`.

- `transaction::isolation` (new): the per-shard hold table (thread-local:
  a TXN writes only on its own shard, #499), gated by one `Cell<usize>`
  load, so a shard with nothing held pays one thread-local load per checked
  write and no atomics. Holds are taken by the connection capture
  (`conn_capture`, which now refuses a key another TXN holds before
  capturing, poisons its TXN, and on an error reply releases the holds it
  created) and by the script leg (`run_local_script`). Released by COMMIT
  (also the killed-snapshot path, which leaked intents before) and by
  `abort_logged` only AFTER the restore's compensating records are
  enqueued.
- Checked on every write path: `command::dispatch` (both runtimes' local
  legs, routed SPSC legs, coordinator legs, MULTI/EXEC bodies, every script
  `redis.call`); the monoio inline SET stands down while anything is held;
  `immediate_serve` / the MULTI BLPOP rewrite; `move_core` / `copy_core`;
  `MQ.*`; FLUSHDB/FLUSHALL in their dispatch arms and SWAPDB in both
  intercepts read every shard's published view (`Published`: one writer,
  whole-word stores; loom model `txn_isolation_view` in
  tests/loom_response_slot.rs). `dispatch_read` executes reads only.
- The TXN's own writes run in an `OwnerScope`; the replica's apply of the
  master stream in a `BypassScope`.
- Blocking wakers leave a held key's waiters parked (a BLMOVE whose
  destination is held too) and re-wake them when the hold is released.
- Eviction samplers and active expiry (whole-key sweep, lazy drain,
  hash-field sweep) skip held keys; an expired held key is reaped after the
  TXN ends, and does not latch the moon#1288 backlog.
- INFO stats: txn_open, txn_oldest_age_ms, txn_held_keys,
  txn_conflicts_refused. docs/guides/transactions.md documents the error.

Evidence (base 2e99254+3 cherry-picked review tests vs this commit, both
runtimes, --shards 1 and 4): tests/txn_isolation_1299.rs 20 tests, 18 red on
base (the 2 `disconnect_releases_*` guards pass vacuously there), 20 green
after on monoio and tokio. The PR #1301 review over-capture shapes: 10/10
lose another client's write on base at --shards 1 (9/10 at 4: `SET 100`
routes to another shard), 0 after. review_w1 ...
an_abort_does_not_overwrite_another_clients_acknowledged_write: red -> green;
its two siblings stay red until WS42 (#1300). Gates green: fmt, clippy
--all-targets on both feature sets, fuzz check, cargo test --lib (6938
monoio / 5996 tokio), txn_abort_durability_1285, txn_multikey_undo_500,
txn_kv_wiring, txn_partial_reject(_monoio), inline_read_txn_visibility_807,
scripts_in_multi_894 (tokio: known replica-order failure),
script_key_routing, eviction_reason_del_run_budget_1294,
active_expiry_backlog_drain_1288.

Deferred: no idle-TXN timeout (maintainer decision); a cross-shard
multi-key write whose one leg meets a held key applies its other legs (the
same non-atomic shape as any failed leg); a race between the FLUSH/SWAPDB
pre-check and a hold taken on another shard during the fan-out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…oon#1299)

`abort_logged` released the transaction's holds with an explicit call after
its last await. A dropped abort future (task cancelled mid-await) would have
left the keys held for the life of the process. The release is now an
`isolation::EndOnDrop` guard taken at the top of `abort_logged`: it still
runs after the restore and the enqueue of its compensating records (end of
the function), and also when the future is dropped mid-await.

Evidence: `transaction::isolation::tests::end_on_drop_releases`;
tests/txn_isolation_1299.rs (incl. `disconnect_releases_{1,4}_shard(s)`)
and txn_abort_durability_1285 green on both runtimes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ut lost on kill -9 (moon#1302)

One Cypher CREATE of 6000 nodes keeps 4096 after kill -9; a TXN with
6000 graph rollback records, or a 2,500-node Cypher CREATE pipelined
with its TXN ABORT, answers MOONERR WAL backpressure; after such an
abort a master restart disagrees with its replica (1905 vs 1 nodes);
a TXN.COMMIT materializing 6000 MQ PUBLISHes keeps 4096. Red on the
WS36 base on both runtimes (graph tests skip on the tokio leg, which
builds without the graph feature); --shards 1 and 4.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…nded, not dropped (moon#1302)

Every graph, MQ, workspace and temporal WAL record went through an
unchecked try_send into the shard's 4096-slot append channel, drained
only on the shard's 1 ms tick; a command emitting more records than the
free slots (a 6000-node Cypher CREATE, a large TXN rollback, a
TXN.COMMIT materializing thousands of MQ PUBLISHes) lost the rest after
answering OK.

New shard::wal_append: a full channel spills the record to the owning
shard thread's overflow queue, and the tick's drain_into appends the
channel's records and then the overflow's, in production order, in one
synchronous stretch. No fixed capacity for one command's records, no
drop, identical durability (same tick, same flush), on-disk WAL-v3
format unchanged. ShardDatabases::{wal_append, try_wal_append_required,
try_wal_append_all} and mq_exec's slice appender all route through it;
a record that still cannot be enqueued (writer gone, or produced off
its shard's thread) is counted in
reclamation_wal_append_channel_dropped_total and logged. The shutdown
paths drain the queues before their final flush. New INFO field
reclamation_wal_append_overflow_total.

TXN.ABORT's checked rollback append (PR #1301) therefore no longer
refuses a burst: MOONERR WAL backpressure / txn_rollback_wal_dropped
remain only for a writer that is gone. append_graph_rollback_wal now
borrows the records and reports the accepted prefix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…L accepted (moon#1302)

abort_logged recorded the graph rollback records in the replication
stream BEFORE the checked WAL append: on a refusal the replica held the
whole rollback while the master's WAL held a prefix, and after a master
restart the two disagreed for good. Append first, then replicate
exactly the accepted prefix, in the same no-await stretch (the EndOnDrop
hold guard from moon#1299 still releases only after the records are
enqueued).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ow answers +OK (moon#1302)

The PR #1301 overflow tests accepted either +OK or MOONERR WAL
backpressure; with no capacity limit on the owning shard's thread the
abort must answer +OK and stay aborted after kill -9.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… the AOF replay clock (moon#1283)

The `touchback` probe of REVIEW-FINAL-P5B re-created as a Rust integration
test (the original rvfb_1277_edges.py is gone): 40 keys in five classes
whose verdict after a restart depends on the clock the log was written
under — keys that expired while the server ran and were rewritten before
their DEL was logged (moon#542: INCR -> 1, SET NX -> new), an RMW logged
while alive whose TTL passes during the downtime, long-TTL and persistent
controls. Every live reply is asserted. The AOF files' mtimes are moved an
hour back (and forward) before the restart.

- moon_1283_touch{back,forward}_probe_s{1,4}
- moon_1283_touchback_probe_after_a_rewrite_s{1,4}: the workload lands in
  a new generation cut by BGREWRITEAOF
- moon_1283_mixed_old_new_log_s{1,4}: a stamp-less first generation
  (MOON_OLD_BIN, or this binary with every MOON.TS stripped) followed by a
  stamped one
- moon_1283_the_log_carries_ts_records
- moon_1283_downgrade_read_s{1,4}: a stamped log replayed by
  MOON_DOWNGRADE_BIN (skips with a note when unset)

Red on the base binaries (2e99254, release-fast), keys wrong:
  monoio s1/s4: touchback 20/40, 17/40; forward 10/40, 10/40;
                after a rewrite 12/60, 18/60; mixed 20/70, 11/70
  tokio  s1/s4: touchback 12/40, 12/40; forward 10/40, 10/40;
                after a rewrite 20/60, 13/60; mixed 16/70, 19/70
(the L/E counts vary run to run: the active expiry reaps some keys, and
logs their DEL, in the few ms they sit expired; the rounds are spread over
the expiry cycle so a regressed replay cannot pass by luck).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… intercept (moon#1283)

Replay side of moon#1283 option (a).

- `replay::pseudo`: the ONE intercept for replay-only `MOON.*` records
  (`MOON.TS`, and now `MOON.COLDCUT` / `MOON.SPILLED` too), called first in
  `DispatchReplayEngine::replay_command` — before anything that could skip
  a data record (decision Q6: clock stamps are observations, not data).
  `classify` is a 5-byte prefix check for every data record. New route
  `ReplayRoute::Marker` (never KV history). `TsRecord` encodes
  `MOON.TS <ms>` on the stack (<= 44 bytes, no allocation).
- `replay::clock`: a pin guard now scopes one replayed file. `observe_log_ts`
  sets the judgment clock to the LAST stamp read (not a running max: a
  parked producer's record carries an older stamp); until a file's first
  stamp the mtime pin rules, so a log without stamps replays exactly as
  before. Stamps outside a pin scope, 0, or past year 9999 are ignored.
- Downgrade: an older binary sends MOON.TS to dispatch, gets "unknown
  command" (ReplayRoute::Unhandled) and replays everything around it;
  pinned by `an_older_binary_sees_an_unknown_command_and_skips_it`.

Tests: replay::pseudo::tests (6), incl. the touchback shape through the
production flat-file replay, a stamp-less prefix + stamped suffix, and
last-not-max.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ON.TS (moon#1283)

Writer side of moon#1283 option (a): the log now says which clock each
record was judged under, so a replay no longer depends on the file mtime
(REVIEW-FINAL-P5B: 27-36/40 keys wrong with the mtime an hour back).

- `AofMessage::{Append,AppendSync}` carry `clock_ms` in memory only, like
  `epoch`. `AofWriterPool::fold_stamp` now returns an `AppendStamp`
  (epoch + this thread's cached clock — the `CachedClock` value the
  mutation was judged with), read in the mutation's synchronous section,
  so a record that parks before its enqueue keeps its own clock. The
  synchronous producers stamp themselves. Stamp-taking APIs accept
  `impl Into<AppendStamp>`; a bare `FoldEpoch` converts with clock 0
  (= unknown: no stamp; barriers, tests).
- `aof::record_ctx::RecordCtx` replaces the writers' `last_db: usize`: one
  rule, `prefix(db, clock_ms, empty)`, yields `MOON.TS <ms>` when the clock
  differs from the last stamp in stream order, then `SELECT <db>` when the
  db differs. Applied by the four batch loops (`inject_record_prefixes`,
  the renamed `inject_select_records`, same #455 filter and moon#1187 fast
  path) and per record by every rewrite and overflow drain. A new
  generation `reset()`s it (db 0, no clock). Stamps are carved out of a
  4 KiB arena: no allocation per stamp.
- Injected stamps are lsn 0 in the framed incr (max_lsn unaffected).
- Every generation head (boot seed, flat-file seed, rewrite) now carries
  `MOON.TS <now>` right after `MOON.COLDCUT` (`write_generation_head_at`
  with 0 = no stamp keeps the byte-exact test fixtures).
- `migrate_aof` copies a `MOON.TS` to every shard's incr instead of
  routing it by its argument's hash.
- size_of::<AofMessage>(): 72 -> 80 bytes (pinned by
  `the_clock_costs_at_most_one_word_per_message`).

Cross-lane touch: one-line `FoldEpoch::INITIAL` -> `AppendStamp::INITIAL`
defaults (and two param types) in handler_monoio/{mod,write}.rs,
handler_sharded/{mod,write}.rs, handler_single.rs, shared.rs, txn_abort.rs,
coordinator.rs; `clock_ms: 0` in test literals.

The per-shard WAL v3 KV log (`--appendonly no`) does not carry stamps yet
(it is not the AOF stream); its replay keeps the mtime judgment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
No target exercised AOF *replay*: `resp_parse*` stop at the RESP layer and
`wal_v3_record` at the WAL record decoder. `aof_incr_replay` drives
arbitrary bytes through the three production readers — the framed
per-shard incr (plus the ordered-entry merge), the multi-part RESP incr,
and the flat `appendonly.aof` (via a temp file, RDB preamble detection
included) — into real databases through `DispatchReplayEngine` and its
pseudo-command intercept. Byte 0 selects the reader and prepends
well-formed `MOON.COLDCUT` / `MOON.TS` / malformed `MOON.TS` records so the
intercept arms are reached without discovery. Invariants: no panic; the
replay clock never leaks out of its pin scope.

- `aof_manifest::shard_replay` is now `pub mod` so the cfg(fuzzing)
  entry `shard_replay::fuzz::{replay_framed, replay_resp}` is reachable.
- Listed in BOTH fuzz.yml matrices (PR + nightly) and in fuzz/Cargo.toml;
  `cargo check --manifest-path fuzz/Cargo.toml --all-targets` is clean.
- No nightly toolchain on this host: a stable smoke of the same contract
  runs as a unit test (`mutated_stamped_logs_never_panic_nor_leak_the_clock`,
  6000 seeded mutations of a stamped log, both readers).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…S (moon#1283)

Lists the replay-only pseudo-commands (SELECT injection, MOON.COLDCUT,
MOON.TS, MOON.SPILLED), what MOON.TS carries and how a replay uses it
(last stamp, mtime fallback until a file's first stamp), and the
compatibility story: no §2 rule-4 layout changes; upgrade replays old logs
as before; a downgraded binary treats MOON.TS as an unknown command, skips
it and replays everything else (verified: moon_1283_downgrade_read_s{1,4}
against the 2e99254 binaries, both runtimes); the WAL v3 KV log does not
carry stamps yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…1283)

The TopLevel writer's `last_db` and the rewrite drains' `db_ctx` are
`RecordCtx`s now; their comments still described the db-only context.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
tests/aof_everysec_kill9_1266.rs acks N SETs under `--appendonly yes
--appendfsync everysec`, SIGKILLs the server 1 ms after the last ack
(MOON_1266_KILL_DELAY_US), restarts on the same --dir and counts acked keys
that did not come back. Shapes: 10,000 unpipelined SETs, 10,000 SETs in
pipelines of 100, and one SET after the writer idled 1.5 s; --shards 1 and
4; 20 reps per cell (MOON_1266_REPS). Every rep's count is printed.

Red on this branch's base (lane-b-ws37-*), 5 reps per cell, lost per rep:
  monoio s1 unpipelined [123,0,104,92,69]   s4 [1,17,3,0,15]
  monoio s1 pipelined   [1118,0,200,10000,0] s4 [100,366,338,300,7557]
  monoio lone-after-idle s1/s4 [1,1,1,1,1]
  tokio  s1 unpipelined [2,67,160,92,109]   s4 [348,184,208,310,364]
  tokio  s1 pipelined   [0,152,201,52,101]  s4 [459,235,469,442,543]
  tokio  lone-after-idle s1/s4 [1,1,1,1,1]

Also: `always` acks only after its fsync (the fsync held open with
MOON_TEST_AOF_SYNC_GATE: no reply while held, +OK once released), and an
everysec writer keeps acking and writing while its fsync is held, with 0
lost to a kill -9 during the held fsync (INFO aof_delayed_fsync >= 1 proves
the hold bit). Those two need the hook and INFO field added with the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…(moon#1266)

The tokio writers kept a batch in their user-space BufWriter until 8 KiB
had accumulated (the moon#1187 tail bound), so a SIGKILL took up to 8 KiB of
acknowledged records per shard with it under everysec/no. A kill -9 does
not touch the kernel page cache: once write(2) has returned, the record
survives. Every non-`always` batch now ends with `flush()`, which also
waits for the write tokio::fs::File may still have in flight on its
blocking pool (its poll_write returns before the write is done).

Cost: one blocking-pool hop per batch instead of one per 8 KiB; the
BufWriter capacity (2 MiB, moon#1226) still sets how many hops a large
batch takes. `always` is unchanged (it already flushed before its fsync).

Evidence: tests/aof_everysec_kill9_1266.rs (tokio numbers in the SUMMARY).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…read (moon#1266)

The everysec fsync ran inline on the AOF writer thread, once a second: for
as long as the disk took, the writer did not drain its channel, so acked
records piled up in process memory (lost to a kill -9) and past 10k
records producers met the moon#769/#838 backpressure. redis runs this
fsync on a background thread (BIO_AOF_FSYNC); moon now does the same.

- fsync_agent.rs: one `aof-fsync-<idx>` thread per writer (spawned by the
  writer, joined when it exits). At the everysec deadline the writer hands
  it a dup of its fd and goes straight back to its channel. The agent
  records the outcome where the inline fsync did (fsync-latency metric,
  record_everysec_fsync_result -> INFO aof_last_fsync_status /
  aof_fsync_failures). A failed fsync is retried at the next deadline even
  with nothing new written. No agent (thread refused) or a failed dup:
  the writer fsyncs inline exactly as before; nothing is dropped.
- fsync_handoff.rs: one atomic word, IDLE -> IN_FLIGHT (writer CAS) ->
  IDLE (agent finish, outcome stored first; or writer abort). At most one
  fsync in flight per writer; a deadline that finds one running is
  postponed and retried on the next wake, counted once per deadline in the
  new INFO field `aof_delayed_fsync`. Unlike redis, the WRITE is never
  postponed, so a slow fsync no longer exposes acked writes to a process
  crash. Loom model: tests/loom_aof_fsync_agent.rs compiles the real file
  (at most one job; every byte written before a claim covered by its
  fsync; a claim sees the settled outcome; a postponed deadline is retried,
  never dropped; an aborted claim frees the state). Reordering `finish`
  (IDLE before the outcome) makes the loom run fail.
- The writer only hands off when something was written since the last
  hand-off (or the last fsync failed): an idle writer no longer fsyncs a
  clean file every second.
- MOON_TEST_AOF_SYNC_GATE=<path> (test-only) holds every AOF data fsync -
  the agent's and `always`'s per-batch one - while <path> exists.
  MOON_TEST_AOF_FSYNC_STALL_MS keeps holding the WRITER at its deadline (it
  now models a writer blocked behind a slow disk), so the moon#769/#838/
  #1272/#1294 suites keep their mechanism.

`appendfsync always` is untouched: its per-batch fsync stays on the writer,
before the batch's acks (group_commit); tests/aof_everysec_kill9_1266.rs
always_acks_only_after_the_fsync_{s1,s4} hold the fsync and see no reply
until it is released.

Cross-ownership: src/command/connection.rs (one INFO line).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…poll, then park (moon#1266)

The monoio writer polled its channel park-free with `wait/16` sleeps:
~3 ms between polls while writing (50 ms floor wait) and 50 ms for the
first record after an idle second (1 s escalated wait). Every acked record
sat in process memory for up to one step, and a kill -9 inside it lost the
record - 10,000 of 10,000 acked pipelined SETs in one base rep.

Now, under everysec/no:
- WARM (the previous receive returned a message): poll every 100 us
  (AOF_WARM_POLL_STEP) for up to 5 ms (AOF_WARM_POLL_SPAN). Producer
  try_sends stay futex-free while writes flow - the reason the loop was
  made park-free (149,718 shard-thread futex wakes per 8 s at p1).
- then, or when the previous receive timed out (COLD): park in
  recv_timeout for the rest of the wait. The first record after an idle
  period costs its producer ONE futex wake and is picked up at once; an
  idle writer wakes only when its wait (<= 1 s) elapses, fewer times than
  the old <= 20/s.
`always` keeps its parked receive.

Unit: poll_recv_tests::a_cold_writer_picks_up_the_first_record_promptly
(median pickup < 10 ms; old step 50 ms), a_warm_writer_polls_with_a_short_step
(median < 1.5 ms; old 3.125 ms), a_warm_writer_parks_after_its_span_and_times_out_on_schedule.
Integration: tests/aof_everysec_kill9_1266.rs (numbers in the SUMMARY).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…r_task/poll.rs (moon#1266)

Pure move of poll_recv / recv_next / the warm-poll constants and their
tests out of the over-cap writer_task.rs (2317 -> 2079 lines, below its
2154-line base). No behavior change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…arm poll step (moon#1266)

A diagnostic knob (10..=50,000 us, read once; default 100 us) for
same-binary A/B runs of the warm poll step's trade-off: the step is the
everysec kill -9 window while writes flow, and ~1/step is the writer's
wake rate. Documented in docs/internal/env-knobs.md, together with the
test-only MOON_TEST_AOF_SYNC_GATE / MOON_TEST_AOF_FSYNC_STALL_MS hooks and
redis's 2 s write-postpone caveat.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ICT opt-in for 0 (moon#1266)

Option 3 narrows the everysec kill -9 window to one warm poll step or one
thread wake-up; it does not close it. On this 4-vCPU host shared with two
other build lanes, 20 reps per cell on the fix still lost acked SETs in 1-2
reps of some cells (a writer descheduled, or a write(2) stalled, for longer
than the 1 ms kill delay), while every base cell lost in its MEDIAN rep.

So the default assertion is the property the fix does deliver: the median
rep loses nothing and at most a quarter of the reps lose anything.
MOON_1266_STRICT=1 asserts 0 in every rep - the bar for moon#1266 1A
(write before the reply, WS46). Per-rep counts are always printed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…le probes for the positional rule (moon#1283)

tests/aof_replay_clock_1283.rs, R2 review of moon#1283:
- downgrade_then_reupgrade(shards, graceful): the newer binary is stopped
  CLEANLY (SHUTDOWN) before the downgrade, as the procedure requires; after
  the re-upgrade, ONE SET and an immediate restart (graceful: *_s1/_s4;
  kill -9: *_then_kill9_s1/_s4), then more writes and a third boot; every
  live key must survive each boot and no AOF rewrite may happen. Red on
  r1b-057598f-monoio (MOON_DOWNGRADE_BIN=lane-b-base-monoio): 351/351 (s1)
  and 86/347 (s4) keys lost after a graceful restart, 338/338 and 86/332
  after kill -9.
- touch_forward_last_tick (NEW-A, the reviewer's q5): SET k 10 PX 600;
  INCR k, kill -9 or graceful stop, mtime moved an hour forward: k absent.
  Red on r1b: 1/8 (s1) and 4/8 (s4) keys resurrected as a persistent 1.
- pure_lifecycle (q1): graceful, kill -9 and two write-less restarts; every
  clean-close marker in every incr is followed by a stamp or EOF, >= 4 per
  log, every key keeps its verdict, and the rule's log line never appears.
  Red on r1b (it writes no marker).
- downgrade_read now stops the newer binary gracefully, so the older binary
  replays (and skips) a clean-close marker; red on r1b for the same reason.
Green on r2fixb-v1-monoio: 19/19, touchback/touchforward/mixed probes
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…1266)

R1 review NIT #10, carried by R2: the `SettleOnUnwind` guard in
fsync_agent.rs (dead stored with Release, then finish(false)) and the
writer's dead check in EverysecSync::dispatch (after try_begin's CAS
Acquire of IDLE, dead.load(Acquire)) were outside the loom model.

tests/loom_aof_fsync_agent.rs gains DyingWorld: the agent takes its first
job and unwinds (guard order as shipped), then its receiver drops; the
writer's two deadlines claim, check `dead`, and send or abort to inline.
Checked under every interleaving: the first claim is sent, the claim after
the death is Inline and sees the failure, no job is left in the dying
agent's still-open queue, the hand-off is never stuck IN_FLIGHT (6.). A
negative control (dead stored AFTER finish) is found by loom:
`(6) a job was sent to the dying agent and is lost in its queue`.

Run: cargo rustc --release --test loom_aof_fsync_agent -- --cfg loom in
/home/user/wt/target-loom (RUSTFLAGS="--cfg loom" breaks hyper-util's
build), then the test binary: 5 passed (loom_dead_agent_is_never_dispatched_to,
loom_dead_flag_after_finish_is_caught - should panic, and the three existing
models). The std smoke variant runs 2,000 rounds under plain cargo test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…opened file's first append (moon#1283)

The record_ctx module doc still said a clock_ms of 0 never emits a MOON.TS;
since 9b8d391 a reopened file's first append is stamped with the writer's
clock (the positional rule's session stamp). Comment only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ed writer's session stamp (moon#1283)

tokio_per_shard_writer_latches_after_torn_write writes byte-exact frames into
a reopened per-shard incr. Since 9b8d391 the writer's first append carries
the session MOON.TS (even with clock_ms 0) before its SELECT 0, so the tear
now falls on the 4th append (was the 3rd), and the file holds the stamp,
SELECT 0 and lsn 1 — and no clean-close marker after the tear (the latch
suppresses it at the cancel). tokio lib persistence: 1050 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ic fold

held_spill_files_are_released_after_the_next_rewrite asserted the held
spill files were still on disk before its manual second BGREWRITEAOF. Since
moon#1289 a database whose files stay held for three orphan sweeps asks the
auto-rewrite monitor for a fold, and with the test's 1 s sweeps that fold
commits inside the 5 s wait: the files are released by a later fold, as the
test intends, just not by the manual one. It failed on the wave-2a tree
("34 at the fold, 0 now") and passed on main cf6fa65.

The precondition now holds when either the files are still there or
INFO cold_held_release_folds_requested > 0; the manual rewrite and its
release assertion run only when files remain; the kill -9 recovery check
of every probe runs in both cases, unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… a TXN write at its start (moon#1289)

Round-2 review N1 (MAJOR). The held-file snapshot's `txn_open` check in
`snapshot_request::request` is one instant on the requesting thread; each
shard starts its part later, at its own tick. A TXN whose BEGIN + SET landed
in between was serialized, and without an AOF a kill -9 restart brought the
aborted write back (reviewer: 12/12 restarts at 4 shards on 057598f).

Mechanism (`persistence::snapshot_request::txn_round`,
`shard::snapshot_txn_guard`):
- `request()` registers the epoch as an automatic round BEFORE it is
  broadcast (`bgsave_start_sharded_announcing`), only for reasons whose
  `waits_for_open_txns()` is true. BGSAVE / SAVE / save rules / SHUTDOWN's
  save never register one and are untouched (WS42, moon#1300).
- Each shard, at the instant it starts its part (`check_auto_save_trigger`,
  before any state, temp file or hold stamp exists), reads its OWN
  thread-local hold table (`isolation::any_held`). A hold abandons the round.
- A shard's file is renamed by its own writer thread, so a shard whose walk
  is done waits before `begin_finalize` until every shard has passed its
  start check; then all publish or all drop. An abandoned part is aborted
  quietly (`SnapshotState::abandon`), COW disarmed,
  `note_snapshot_finished(false)` (no held file released), state dropped (the
  stream writer's cancel path removes the temp file). A walk in progress
  stops early (`txn_round::is_abandoned`, one load per tick).
- Fan-in: `bgsave_shard_abandoned()`; a round that ends abandoned records no
  outcome (LASTSAVE, rdb_last_bgsave_status, dirty counter untouched) unless a
  shard genuinely failed.
- The gate slot the request took is given back (`SnapshotGate::give_back`),
  so a later sweep asks again; counted once per round in the new INFO field
  `cold_held_release_snapshots_abandoned_txn` (reset by CONFIG RESETSTAT).
  Logged once at warn by the shard that found the hold.
- The early `txn_open` check stays as the cheap pre-filter.

Why this is complete: a TXN writes only on its connection's shard (#499),
and its hold is taken BEFORE the write is dispatched
(`conn_capture::capture_conn_write`; the script leg holds inside the same
synchronous run) and released only after an abort's restore is applied
(`isolation::txn_end`). So a shard with no hold at its start has no
uncommitted write in memory. A TXN write that lands after that shard's start
reaches the keyspace through `command::dispatch` (both runtimes' TXN legs),
a script's `redis.call`, or the abort's `kv_compensation::undo_one`, and each
captures the key's epoch-start image first, first capture wins
(`snapshot_cow`): the file holds the pre-TXN value. Verified in code and
pinned by `snapshot::txn_after_start_tests` (the handlers' exact
capture_conn_write + dispatch sequence, TXN left open through the publish,
and an abort landing mid-walk).

No new atomic state machine: the round is a parking_lot::Mutex<Option<Round>>
plus one published epoch (a plain value), no loom model needed.

Evidence (Linux container, not the merge bar):
- tests/held_release_txn_race_1289.rs (new; hook
  MOON_TEST_SNAPSHOT_START_HOLD_FILE holds every shard at its start so the
  TXN lands in the window by construction):
  - `a_txn_between_the_request_and_the_start_is_not_saved_{1_shard,4_shards}`:
    red on 057598f + the hook only (r2fixc-hookonly-{monoio,tokio}): restart
    answers k=aborted new=inserted; plain 057598f fails the "start is held"
    precondition. Green on r2fixc-n1-{monoio,tokio}: k=original, new=nil,
    abandoned=1, no shard file renamed, no temp file, LASTSAVE unmoved,
    status ok, held files still on disk.
  - `an_abandoned_round_is_retried_once_the_txn_ends`: green (the retried
    snapshot released the files ~0.8 s after the ABORT).
  - `a_busy_txn_workload_never_reaches_the_automatic_snapshot_{1,4}`: the
    reviewer's txnrace in-suite; s1 red on r1b-057598f-monoio, both green on
    the fix (s4: deferred=45 abandoned=15 before a clean publish).
- Unit: txn_round (6), snapshot_request give-back, fan-in
  `an_abandoned_round_ends_without_an_outcome`, txn_after_start_tests (2).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… sees as past; the master sends DEL (moon#1286)

Round-2 review N2. R1's F4 (redis `checkAlreadyExpired`: an absolute
deadline already past deletes the key at once) also ran on a replica
applying its master's stream. A lagging replica therefore deleted keys its
master kept and published `del` for them (reviewer's replag.sh: replica
SIGSTOPped over `PEXPIREAT k now+100; PERSIST k` ended at DBSIZE 1 with
del:b, del:c; redis 7.2.7's replica keeps all three).

- `command::key::deadline_already_past(db, when)`: the F4 test, false while
  `applying_master_stream()`. Used by the four F4 arms (EXPIREAT, PEXPIREAT,
  GETEX EXAT/PXAT, RESTORE ... ABSTTL). On the master stream the deadline is
  stored like any other, as redis's replica does.
- The master now says what it did, as redis's `rewriteClientCommandVector`
  does: `replication::effect_rewrite` propagates EXPIREAT / PEXPIREAT that
  answered :1 with a deadline <= now, and GETEX EXAT/PXAT past that answered
  the value, as `DEL key`; RESTORE ... ABSTTL past as `DEL key` with REPLACE
  and nothing without (it wrote nothing). Without this half the replica would
  now keep an invisible key forever (a replica runs no expiry of its own);
  before, it only stayed consistent because it deleted by its own clock.
- AOF replay keeps F4: `applying_master_stream()` is false there and the
  database clock is pinned to the log's time (moon#1277), so a replayed
  deadline is judged exactly as it was live - deleting reproduces the live
  state with no post-boot `expired` reap. redis skips while loading because
  its clock is not pinned. New logs carry the DEL anyway.

Not changed (pre-existing, not a one-line fix): the replica still judges
expiry by its own clock on master-stream LOOKUPS, so in the lag repro the
following PERSIST / PEXPIRE misses a key the replica sees as expired - the
keys now stay resident (DBSIZE matches) but hidden until the master next
writes them. redis's `expireIfNeeded` answers "live" to the master client;
moon would need that in every `is_expired_at` lookup reached from apply.

Evidence (Linux container, not the merge bar):
- tests/replica_past_deadline_1286.rs (new, monoio: master PSYNC):
  `a_lagging_replica_does_not_delete_what_its_master_kept` red 3/3 on
  r1b-057598f-monoio (`del` published for a, b, c), green 2/2 on
  r2fixc-n2-monoio; `a_past_deadline_on_the_master_deletes_on_the_replica_too`
  green on both (guards the DEL half: skipping on the replica without it
  leaves d, e resident there).
- Unit: effect_rewrite `a_past_absolute_deadline_propagates_as_del` (7 DEL
  forms, 1 skip, 6 verbatim controls); expire_past_tests
  `a_replica_applying_its_master_stream_stores_a_past_deadline` (4 commands
  under the master-stream scope keep the key with the deadline).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…s (moon#1286)

Round-2 review N4. R1's F3 counts a write that lands on an expired HOT key
(redis `lookupKeyWrite` -> `expireIfNeeded`). A key only the cold tier held
went through `Database::set_recording`'s miss arm, which inserted the new
hot value and left the dead cold entry uncounted (reviewer: 1,500 spilled
keys with PX 6000, SET after expiry: delta 186 - only the hot ones the
active cycle had reaped - where redis counts 1,500).

The cold entry's deadline is in the in-RAM index (`ColdLocation::ttl_ms`),
so no I/O is needed: on `Inserted`, when the cold index is non-empty, one
index lookup (`ColdIndex::expired_at`, the same `now > ttl` judgement as the
expiry sweep and `remove_counting_cold`). An expired entry is counted and
removed at once, so no later sweep or read can count it again; a live cold
shadow is left exactly as before (the task #56 rule). The lookup runs only
for a NEW hot key while something is spilled; no allocation. The counter
still skips replay and a replica's master stream (`counts_expiry`).

Evidence: unit `a_write_over_an_expired_cold_only_key_counts_it_once` red
on the parent (counted 0, want 1), green after; it also checks the
overwrite and a live shadow count nothing and the live shadow stays.
Integration below (expired_keys_parity_1286, info_expired_keys_1286) at the
gates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… and what it can starve (moon#1289)

Round-2 review N3. "It never captures a transaction's uncommitted writes"
was an overclaim before the N1 fix. It now states the three mechanisms that
make it true for the automatic held-file snapshot (deferred while a TXN is
open; abandoned whole if a shard holds a TXN write at its start; later TXN
writes saved at their pre-transaction value by copy-on-write), that BGSAVE,
SAVE, the save rules and SHUTDOWN's save are not covered (moon#1300), and
the starvation cost: an open TXN or unbroken TXN traffic keeps held files on
disk and SWAPDB refused, with no timeout - visible in
cold_held_release_snapshots_deferred_txn / _abandoned_txn while
cold_held_files_stale_databases stays above 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…hen sampling finds no TTL key (R2 item 5)

The tokio lib-suite failure of
`snapshot::eviction_capture_tests::a_failed_save_frees_a_held_victim_off_the_shard_thread`
("evicts to budget: OOM command not allowed ...") is not shared process
state. Its fixture holds ONE TTL key among 301 keys in 8 DashTable segments
under volatile-lru; `sample_victim` draws random segments and gives up after
8 x maxmemory_samples = 40 draws, so it misses the only volatile key with
probability (7/8)^40 = 0.48 %. Measured: 14 misses in 3,000 runs of the
fixture on ONE thread with nothing else running (a temporary stress test,
not committed). Three tests share the fixture, ~1.4 % per suite run on
either runtime; "passes alone 3/3" is the same odds.

It is a product bug, not a test bug: with few TTL keys among many segments a
volatile-* policy answered OOM while a volatile key existed (redis samples
its expires dict and cannot miss one). `select_victim` now falls back to the
exact nearest-deadline victim (`find_victim_volatile_ttl`, the deadline
index, O(log n), skips a TXN-held key) when volatile-lru / -lfu / -random
sampling found no candidate. Only the former OOM path changes.

Isolation hardening in the same place: those eviction assertions have no
budget slack, and `footprint_correction()` IS process-global (the
`correction_round_trips_a_published_ratio` test stores 2.5 in it; any
in-process shard-0 chore republishes it). A test-only per-thread pin
(`admin::footprint::pin_footprint_correction_for_test`, RAII guard, cfg(test)
only) holds it neutral for the three big-victim tests and the new test. No
assertion loosened, nothing skipped.

Evidence (Linux container):
- New `storage::eviction::volatile_fallback_tests::a_lone_volatile_key_is_always_found`
  (2,000 plain keys + 1 TTL key, >= 16 segments, 50 rounds x 3 policies):
  red with the fallback reverted ("volatile-lru round 1: OOM with a volatile
  key present"), green with it.
- eviction_capture, footprint and storage::eviction unit tests: 96/96.
- Full tokio lib suite x3 at the gates (see SUMMARY).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
… (clippy, moon#1289)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ffic (moon#1286)

Follow-up to be90b30 (review N4), found by the full monoio lib gate:
`cold_del_rewrite_tests::a_deleted_cold_key_whose_ttl_passed_needs_no_del_and_stays_dead`
failed its "both keys cold" fixture. During a log replay the cold index is
attached before the log, and the replayed SET of a key IS the write whose
spill the cold slot holds (a later MOON.SPILLED marker names it - the task
#56 rule in the Updated arm). The new Inserted-arm reap removed that slot
when its TTL had passed on the unpinned test clock.

The reap (and its count) now runs only where `counts_expiry()` holds: not
during a replay, not while a replica applies its master's stream (it leaves
cold entries to its master, like hot ones). The counter already had that
gate; the removal now shares it.

Evidence: cold_del_rewrite_tests, cold_index, replay, expiration and the N4
unit test: 236/236 on monoio.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
, #1286, #1289, volatile eviction)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…instead of trusting the request counter

Round-3 review NIT on f6033bd: INFO cold_held_release_folds_requested
only shows a fold was dispatched, not that it committed before the files
were unlinked, so a moon#1231-class premature unlink would still pass.

The promote scenario now samples every 100 ms during its orphan-sweep wait
and, at the first sample with no spill file left, asserts that the AOF
generation (the sum of every shard's base sequence) already advanced past
the one the first fold left, i.e. a later fold COMMITTED before the files
went. A flat legacy appendonly.aof (tokio --shards 1) has no base sequence;
there it falls back to the fold counter. The check now guards all three
promote cases, not only the second-rewrite one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…stead of 'drained and synced' (moon#1283)

Round-3 review R3-1: the downgrade procedure keys on the log line 'AOF
writers drained and synced', but a writer with a latched write error skips
its MOON.TS CLOSE marker and still exits normally, and a panicked writer is
joined too, so the line was logged over a stop the next boot reads as a
crash (repro: ulimit -f until EFBIG, then SHUTDOWN: the line once, 0
markers).

Every successful marker append now counts (writer_stop::note_close_marker);
stop_writers compares the markers written during the stop with the writers
it joined and, when any is missing, logs a warning 'AOF writers stopped ...,
N of M WITHOUT their clean-close marker' instead of the drained line.
STORAGE-FORMAT §3.3 and the production guide name the new line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…ot for every spill file (moon#1231, moon#1289)

held_spill_files_are_released_after_the_next_rewrite failed on the wave-2a
tokio build ("the second rewrite committed but the sweep released no held
spill file (N before, N after)") and passed on main. No product bug: the
test's proxy for "files are still held" was "some spill file is still on
disk", and since moon#1289 that proxy is wrong on tokio.

Mechanism, measured with a driver that replays the scenario and samples
INFO every 500 ms (r2-81fda0c, --shards 4):
- moon#1289's automatic fold commits ~3.5 s after the touches (three 1 s
  sweeps stale + the monitor's 1 s tick), inside the test's 5 s wait, and
  the sweep after it releases EVERY held file: cold_files_pending_unlink
  13 -> 0, AOF generation 8 -> 12.
- A tokio GET can serve a cold key without promoting it (the read-only
  get_cold_value path): 33-38 of the 200 probes stay cold (cold_keys:38
  after a GET of each; again at the restart), so their spill files stay
  REFERENCED and the heap-file count never reaches 0. On monoio every GET
  promotes (cold_keys:0), the count reaches 0, and f6033bd/b55c304's
  auto-fold branch took the case.
- So on tokio the test ran the manual second rewrite with nothing held
  (cold_files_pending_unlink:0) and asserted fewer files: 22 -> 22.
- Main (lane-b-base) has no automatic fold: the files are still held at
  the manual rewrite, which releases them.

The case now asserts what it means:
- Every spill file on disk right after the fold is covered by the hold;
  every 100 ms of the wait, ANY of them that is gone must postdate a
  committed later generation (listing first, generation second). The
  b55c304 check ran only once no file at all was left, which tokio never
  reaches, so there it checked nothing.
- The second rewrite runs iff INFO cold_files_pending_unlink > 0 after
  the wait; it must commit a NEW generation (resent if an automatic fold
  holds the flag; rewrite_and_wait could not tell, every shard already had
  a compacted base), and within 10 s the sweeps must bring
  cold_files_pending_unlink to 0 and the file count down.
- Otherwise the automatic fold released them: at least one at-fold file
  must be gone (after a later commit), or the case proved nothing.

Evidence (tokio harness, --include-ignored --test-threads=1 held_spill):
- before: lane-b-base-tokio 8/8 pass; r2-81fda0c-tokio 0/9 (N/N, N 14-21)
- after: r2-81fda0c-tokio 8/8, r2-81fda0c-monoio 8/8, r3fixd-v1-tokio
  10/10, r3fixd-v1-monoio 10/10; lane-b-base-{tokio 3/3, monoio 2/2} via
  the manual path ("36 held after the wait, 0 held after the second
  BGREWRITEAOF"); MOON_TEST_COLD_DEL_SHARDS=1 promote cases green on both.
- mutants (tokio): hold released at once -> "7 of the 32 spill files on
  disk at the fold ... were unlinked before any later fold committed (AOF
  generation 8 -> 8)" (the old test blamed it on the second rewrite);
  hold never released -> "did not release every held spill file within
  10s (33 files and 16 held before; 33 files and 16 held after)".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…n nothing is spilled (R3 fix-e, no issue filed)

Investigating the wave-2a write-path CPU regression (+10.6% CPU/op claimed for
SET p16 c50, --shards 1, no AOF) found no wave-2a cost worth removing, and one
pre-existing cost that is larger than the whole claimed regression.

Mechanism: with disk offload on (the default), every hot-key overwrite runs
`set_recording` -> `ColdIndex::remove(key)` to retire a stale cold shadow.
With nothing spilled, that call still hashed the key (`scan_h48`, xxh64) and
range-probed an empty BTreeMap before answering `false`. It now returns
`false` straight away when both `map` and `older_copies` are empty, which is
exactly the state in which the old walk had no side effect
(`release_older_copies_of` returned on an empty `older_copies`, `remove_raw`
found nothing). The full path moves unchanged into `remove_present`.

Evidence (callgrind, release-with-debug, MOON_NO_URING=1, --shards 1
--appendonly no --save "" --maxmemory 0, 300k SET p16 c10 -r 100000 -d 64 on a
pre-warmed keyspace, shard thread only):
  main cf6fa65       2795.9 Ir/op
  wave 2a 7271e74    2806.9 Ir/op   (+11.1, +0.4%)
  this commit        2605.9 Ir/op   (-201, -7.2% vs 7271e74; -6.8% vs main)
  set_recording inclusive 1052.6 -> 844.7 Ir/op; ColdIndex::remove (206
  Ir/op inclusive, xxh64 ~94 self) no longer appears.

Wave-2a delta, by function (+11.1 Ir/op total): set_recording +7.7 (moon#1286
expired-overwrite compare and the `cached_now_ms` load), try_inline_dispatch_loop
+5.5 self (moon#1299 `any_held()` gate on the inline SET), the moon#1299 exit
wrapper / `&mut ConnectionState` indirection ~+5, offset by -5.7 elsewhere.
check_write never runs on this path (inline SET), and no io_uring event count
moved: 125,206 submit_req/complete per 1M SET on both.

CPU/op (cpu.sh workload, 1M SET p16 c50, io_uring monoio, unpinned, 12 reps
rotating order, release profile for all arms):
  main 985 | 7271e74 980 (-0.5%, 95% CI [-5.0%, +7.1%]) | this 960 (-2.5%, CI [-5.5%, +1.6%])
cpu.sh as written (5 reps, fixed order): main 920,980,1060,860,880 vs this
920,970,900,910,920 (medians 920 / 920); main vs bda76c1 re-run: medians 910 /
930 (+2.2%). tokio (8 reps rotating): main 2435 | bda76c1 2355 | this 2415.

Test: cold_index::tests::remove_on_an_empty_index_is_a_no_op_and_still_releases_older_copies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
…rly return)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
@qodo-code-review

Copy link
Copy Markdown

ⓘ Qodo reviews are paused because the subscription is no longer active. Ask your workspace admin to reactivate the subscription to resume reviews. Manage billing

@coderabbitai

coderabbitai Bot commented Oct 1, 2026

Copy link
Copy Markdown

Important

Review skipped

Too many files!

This PR contains 166 files, which is 66 over the limit of 100.

To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch.

Upgrade to a paid plan to raise the limit.

This review couldn't start because sufficient usage credits or metered capacity aren't available. Add credits or update usage-based reviews in the billing tab, then retry.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 6f323270-6aff-43d3-9775-05f0f85feac3

📥 Commits

Reviewing files that changed from the base of the PR and between cf6fa65 and fa3f751.

📒 Files selected for processing (166)
  • .add/milestones/v0-9-2-perf-review/TEAM-RULES.md
  • .add/milestones/v0-9-2-perf-review/plans/R1-fix-a/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/R1-fix-b/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/R1-fix-c/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/R2-fix-a/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/R2-fix-b/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/R2-fix-c/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/R3-fix-d/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/R3-fix-e/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/WAVE2-PLAN.md
  • .add/milestones/v0-9-2-perf-review/plans/WS36-txn-isolation/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/WS37-aof-ts/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/WS38-parity/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/WS39-held-files/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/WS40-fsync-agent/SUMMARY.md
  • .add/milestones/v0-9-2-perf-review/plans/WS41-graph-wal/SUMMARY.md
  • .github/workflows/fuzz.yml
  • CHANGELOG.md
  • README.md
  • docs/PRODUCTION-CONTRACT.md
  • docs/STORAGE-FORMAT-V1.md
  • docs/guides/persistence.md
  • docs/guides/transactions.md
  • docs/internal/env-knobs.md
  • docs/production-guide.md
  • fuzz/Cargo.toml
  • fuzz/fuzz_targets/aof_incr_replay.rs
  • scripts/test-commands.sh
  • scripts/test-consistency.sh
  • src/acl/command_rules.rs
  • src/acl/io.rs
  • src/acl/mod.rs
  • src/acl/rules.rs
  • src/acl/subcommand.rs
  • src/acl/table.rs
  • src/admin/footprint.rs
  • src/admin/metrics_setup/mod.rs
  • src/blocking/pop_log.rs
  • src/blocking/stream_wake.rs
  • src/blocking/wakeup.rs
  • src/command/config.rs
  • src/command/connection.rs
  • src/command/dump_restore.rs
  • src/command/expire_past_tests.rs
  • src/command/info_reclamation.rs
  • src/command/key.rs
  • src/command/keyspace/move_cmd.rs
  • src/command/mod.rs
  • src/command/persistence.rs
  • src/command/string/string_read.rs
  • src/persistence/aof/auto_rewrite.rs
  • src/persistence/aof/fsync_agent.rs
  • src/persistence/aof/fsync_handoff.rs
  • src/persistence/aof/group_commit.rs
  • src/persistence/aof/mod.rs
  • src/persistence/aof/pool.rs
  • src/persistence/aof/record_ctx.rs
  • src/persistence/aof/refusal.rs
  • src/persistence/aof/rewrite.rs
  • src/persistence/aof/rewrite/flat_file_fold_tests.rs
  • src/persistence/aof/rewrite/rewrite_stall_1158_tests.rs
  • src/persistence/aof/rewrite_overflow.rs
  • src/persistence/aof/runtime_fsync.rs
  • src/persistence/aof/writer_stop.rs
  • src/persistence/aof/writer_task.rs
  • src/persistence/aof/writer_task/close.rs
  • src/persistence/aof/writer_task/poll.rs
  • src/persistence/aof_manifest/mod.rs
  • src/persistence/aof_manifest/shard_replay.rs
  • src/persistence/aof_manifest/shard_replay_fuzz.rs
  • src/persistence/cold_records.rs
  • src/persistence/migrate_aof.rs
  • src/persistence/mod.rs
  • src/persistence/replay.rs
  • src/persistence/replay/clock.rs
  • src/persistence/replay/clock_tests.rs
  • src/persistence/replay/log_segment.rs
  • src/persistence/replay/pseudo.rs
  • src/persistence/replay/pseudo_tests.rs
  • src/persistence/replay/scope.rs
  • src/persistence/snapshot.rs
  • src/persistence/snapshot/abandon.rs
  • src/persistence/snapshot/eviction_capture_tests.rs
  • src/persistence/snapshot/txn_after_start_tests.rs
  • src/persistence/snapshot_request.rs
  • src/persistence/snapshot_request/txn_round.rs
  • src/replication/apply.rs
  • src/replication/effect_rewrite.rs
  • src/scripting/bridge/redis_call.rs
  • src/scripting/bridge/txn_capture.rs
  • src/server/conn/blocking.rs
  • src/server/conn/blocking_txn.rs
  • src/server/conn/handler_monoio/dispatch.rs
  • src/server/conn/handler_monoio/exit.rs
  • src/server/conn/handler_monoio/mod.rs
  • src/server/conn/handler_monoio/txn.rs
  • src/server/conn/handler_monoio/write.rs
  • src/server/conn/handler_sharded/dispatch.rs
  • src/server/conn/handler_sharded/exit.rs
  • src/server/conn/handler_sharded/mod.rs
  • src/server/conn/handler_sharded/pubsub.rs
  • src/server/conn/handler_sharded/txn.rs
  • src/server/conn/handler_sharded/write.rs
  • src/server/conn/handler_single.rs
  • src/server/conn/shared.rs
  • src/server/conn/single_aof_log.rs
  • src/server/conn/txn_abort.rs
  • src/server/conn/txn_script_undo.rs
  • src/server/expiration.rs
  • src/server/expiration_expired_keys_tests.rs
  • src/shard/coordinator.rs
  • src/shard/coordinator/refused_leg_tests.rs
  • src/shard/coordinator/swapdb_fold_tests.rs
  • src/shard/event_loop.rs
  • src/shard/held_release_tick.rs
  • src/shard/mod.rs
  • src/shard/mq_exec.rs
  • src/shard/persistence_tick.rs
  • src/shard/shared_databases.rs
  • src/shard/snapshot_txn_guard.rs
  • src/shard/spsc_handler.rs
  • src/shard/spsc_handler/aof_admission_tests.rs
  • src/shard/test_hooks.rs
  • src/shard/timers.rs
  • src/shard/wal_append.rs
  • src/storage/db/accessors.rs
  • src/storage/db/cold_promote.rs
  • src/storage/db/kv_ops.rs
  • src/storage/eviction.rs
  • src/storage/eviction/volatile_fallback_tests.rs
  • src/storage/tiered/cold_index.rs
  • src/storage/tiered/cold_index/tests.rs
  • src/storage/tiered/cold_reclaim.rs
  • src/storage/tiered/cold_reclaim_tests.rs
  • src/storage/tiered/held_release.rs
  • src/storage/tiered/mod.rs
  • src/storage/tiered/snapshot_hold.rs
  • src/transaction/abort.rs
  • src/transaction/conn_capture.rs
  • src/transaction/isolation.rs
  • src/transaction/kv_mvcc.rs
  • src/transaction/mod.rs
  • src/transaction/undo_log.rs
  • src/workspace/mod.rs
  • tests/acl_rule_order_1296.rs
  • tests/aof_everysec_kill9_1266.rs
  • tests/aof_fsync_stall_r1.rs
  • tests/aof_replay_clock_1283.rs
  • tests/aof_select_after_restart_r1.rs
  • tests/cold_held_files_release_1289.rs
  • tests/crash_recovery_cold_del_rewrite.rs
  • tests/expired_keys_parity_1286.rs
  • tests/graph_wal_append_1302.rs
  • tests/held_release_txn_open_1289.rs
  • tests/held_release_txn_race_1289.rs
  • tests/info_expired_keys_1286.rs
  • tests/loom_aof_fsync_agent.rs
  • tests/loom_response_slot.rs
  • tests/replica_past_deadline_1286.rs
  • tests/review_r1_txn_isolation_1299.rs
  • tests/review_w1_txn_abort_no_aof_snapshot_1285.rs
  • tests/txn_abort_durability_1285.rs
  • tests/txn_close_after_epilogue_1299.rs
  • tests/txn_exit_epilogue_1299.rs
  • tests/txn_isolation_1299.rs
  • tests/wal_group_commit.rs

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@TinDang97
TinDang97 merged commit 9ffa7f5 into main Oct 4, 2026
14 checks passed
@TinDang97
TinDang97 deleted the claude/gifted-mendel-e9wiz5 branch October 4, 2026 16:30
TinDang97 pushed a commit that referenced this pull request Oct 4, 2026
… --onto main

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D81NZwUjEnWCaZekSCB3Uu
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants