Skip to content

feat(bulk): release export/submit leases on graceful shutdown (#1531) - #1538

Merged
smunini merged 2 commits into
mainfrom
feat/1531-release-leases-on-shutdown
Sep 27, 2026
Merged

smunini merged 2 commits into
mainfrom
feat/1531-release-leases-on-shutdown

Conversation

@smunini

@smunini smunini commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Closes #1531.

Summary

On a graceful shutdown (SIGINT/SIGTERM), the bulk export and submit workers now stop claiming, stop their current job at a safe boundary, and release the lease. Another instance can claim the job at once instead of waiting out HFS_BULK_{EXPORT,SUBMIT}_LEASE_DURATION.

Export release (SQLite, PostgreSQL)

One fenced transaction (worker_id + fencing_token, status in-progress):

release now returns bool. A zombie release after a reclaim is a no-op.

Submit release (SQLite, PostgreSQL, MongoDB, S3, composite wrappers)

Audited as the issue asked. A manifest released to pending ends up in the same state as one whose lease lapsed, and that path is already safe:

So release only gains a bool return. The S3 version no longer swallows storage errors.

Workers

  • DefaultExportWorker::with_shutdown: stops between batches, or mid-read/_typeFilter search where no part is open, then releases.
  • DefaultSubmitWorker::with_shutdown: a bridge task trips the existing cooperative cancel (bulk-submit: Abort cannot stop an in-flight manifest, and a failed abort is invisible #968). The wind-down paths release when the stop came from shutdown and the lease is still held. A concurrent abort wins: the release is fenced on processing.

hfs wiring

  • New crates/hfs/src/worker_shutdown.rs tracks the worker loops (TaskTracker + CancellationToken).
  • The shutdown signal cancels the token. After the HTTP drain, serve() waits up to HFS_WORKER_SHUTDOWN_TIMEOUT (default 20 s), then records the audit shutdown and flushes audit/OTLP. Past the deadline, leases lapse as before.
  • Claims are never raced against shutdown, so a claim cannot commit and then be dropped.

Shutdown signal

Hooked into helios_observability::shutdown::signal() (SIGINT and SIGTERM, #907), so a rolling restart's SIGTERM releases leases too. The worker drain runs after the HTTP drain and before the audit/OTLP flush, including on a serve error.

Tests

  • SQLite and Postgres export: claim, partial run, release, reclaim under a cap of 1 gives the same attempt count and no leftover rows. Zombie release and release-after-finish are no-ops.
  • Shared tests/bulk_submit/release_contract.rs, run on SQLite, Postgres, MongoDB and S3: release re-queues; zombie release, double release and release-after-abort are no-ops.
  • Worker-level shutdown tests for export and submit (inline and independent file scheduling): the next worker completes with every resource.
  • worker_shutdown unit tests.

Verified locally: persistence lib tests (sqlite, s3), sqlite_tests and the SQLite bulk-submit suites, the Postgres bulk/export and Mongo bulk_submit filtered suites, hfs tests, and clippy with the CI flags.

🤖 Generated with Claude Code

smunini and others added 2 commits September 26, 2026 12:55
On Ctrl-C the bulk export and submit workers now stop claiming, stop
their job at a safe boundary, and release the lease, so another instance
can claim it at once instead of waiting out the lease duration.

- Export `release` (SQLite, PostgreSQL): in one fenced transaction the job
  returns to `accepted`, its attempt is refunded, and its progress and file
  rows are wiped. Returns whether the lease was still held; a zombie
  release is a no-op.
- Submit `release` (SQLite, PostgreSQL, MongoDB, S3): returns whether it
  took effect. Re-queuing to `pending` matches the existing lapsed-lease
  reclaim path, which is already safe (idempotent re-walk, no attempt
  counter, artifacts published only with the terminal state).
- Workers take a `CancellationToken` via `with_shutdown`. Export stops
  between batches or mid-read; submit reuses the cooperative abort stop.
- hfs tracks the worker loops, cancels them on the shutdown signal, and
  waits up to HFS_WORKER_SHUTDOWN_TIMEOUT (default 20s) after the HTTP
  drain and before the audit/OTLP flush.
- Contract tests per backend, plus worker-level shutdown tests.

Co-Authored-By: Claude <noreply@anthropic.com>
…ses-on-shutdown

# Conflicts:
#	crates/hfs/src/main.rs
#	crates/persistence/src/backends/sqlite/bulk_export.rs
#	crates/persistence/tests/mongodb_tests.rs
#	crates/persistence/tests/postgres_tests.rs
@smunini
smunini merged commit 0eb99c2 into main Sep 27, 2026
18 checks passed
@smunini
smunini deleted the feat/1531-release-leases-on-shutdown branch September 27, 2026 12:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Release bulk export/submit leases on graceful shutdown instead of letting them lapse

1 participant