Repository navigation
Bulk Export
#104
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Introduction
If authentication and authorization decide who may access healthcare data, bulk data export decides how much of it can leave the building at one time - and at what pace, and through what door. It is the API that population health platforms, payer-provider data exchanges, registry submissions, research extracts, and AI training pipelines all converge on. CRUD and search are the bread and butter of FHIR, but the moment a workload needs every Observation for every patient in a cohort, it stops looking like a request/response problem and starts looking like a data engineering problem.
This document shares my thoughts on how to approach Bulk Data Export for the Helios FHIR Server. Like the persistence layer discussion and the authentication and authorization discussion, this is an architectural strategy document rather than a comprehensive specification. It explains the motivating direction, the key building blocks, and the Rust trait designs that will shape the export subsystem.
Who should read this? Anyone with an interest in FHIR bulk data interoperability, healthcare analytics infrastructure, or the operational realities of running long-lived, asynchronous jobs alongside a high-throughput FHIR API. Feedback is very much welcome - this is open source, developed in the open, and your perspective matters.
A note on scope: this document covers export only - the FHIR Bulk Data Access IG, specifically the
$exportfamily of operations. The companion problem of bulk submit - the inverse direction, taking large NDJSON payloads back into the server - is a separate concern with separate trade-offs. The Argonaut Project's current draft of$bulk-submitis being worked through here; we will publish a separate discussion document for ingestion once that draft stabilizes.The Lay of the Land: What the Bulk Data Access IG Says
The Bulk Data Access IG defines an asynchronous, manifest-based, NDJSON-over-HTTPS pattern for exporting large volumes of FHIR data. The shape of every export, regardless of scope, is the same:
$exportoperation. The server responds immediately with202 Acceptedand aContent-Locationheader pointing to a status URL.202. When the export is complete, the server returns200 OKwith a JSON manifest describing every output file.application/fhir+ndjson- one resource per line, no Bundle wrapper.DELETEon the status URL) when finished, signaling that the server may reclaim the output files.Three flavors of export sit on top of this same pattern:
[base]/$export- everything the caller is permitted to see across the entire server.[base]/Patient/$export- every resource in every patient's compartment that the caller is permitted to see.[base]/Group/[id]/$export- every resource in the compartment of every patient who is a member of the named Group.The kick-off request accepts a substantial parameter surface. Some of the headline parameters that any compliant implementation must understand:
_type_sinceinstant. Only resources whosemeta.lastUpdatedis at or after this point._untilinstant. Only resources whosemeta.lastUpdatedis at or before this point._typeFilterMedicationRequest?status=active). May be repeated._outputFormatapplication/fhir+ndjsonis required; abbreviatedapplication/ndjsonandndjsonmust also be accepted._elementsSUBSETTEDtag.includeAssociatedDataLatestProvenanceResources).organizeOutputByallowPartialManifestslink[]pagination before all output is finished.patient(POST only)Patientreferences.The manifest returned at completion has a fixed schema:
{ "transactionTime": "2026-05-11T00:00:00Z", "request": "https://fhir.example.org/Group/cohort-1/$export?_type=Patient,Observation", "requiresAccessToken": true, "output": [ { "type": "Patient", "url": "https://files.example.org/exports/abc/Patient-001.ndjson", "count": 12500 }, { "type": "Observation", "url": "https://files.example.org/exports/abc/Observation-001.ndjson", "count": 980321 } ], "deleted": [], "error": [], "link": [] }Three fields are non-obvious and worth highlighting up front:
transactionTimeis the server's frozen wall-clock at the moment the export was started. Every resource in the output must reflect server state as of that instant. This is the anchor that lets clients implement incremental sync correctly using_sinceon the next run.requiresAccessTokenis a hint to the client about how to fetch the output files. Iftrue, the URLs require the sameAuthorization: Bearer ...token that authorized the kick-off. Iffalse, the URLs are pre-signed (or otherwise capability-based) and the client SHALL NOT send a token. The decision is the server's; both modes are valid.erroris not a status indicator. It is a list of NDJSON files containing FHIROperationOutcomeresources, one per line, describing per-resource-type failures that did not cause the entire job to fail. An export with a populatederrorarray still finished200 OKfrom a workflow perspective.Authorization for bulk operations sits squarely on SMART Backend Services, with scopes of the form
system/Patient.rs,system/*.rs, and similar. That story is told in detail in Discussion #45 and shipped today; we will not re-derive it here. What this document does assume from that work is that the auth layer produces aRequestContextcontaining a validatedPrincipalwith aScopeSet, and that this context flows into every export handler intact.The Essential Flow
In words and then in pictures.
Kick-off. The client sends
GET /Group/cohort-1/$export?_type=Patient,Observation&_since=2026-01-01T00:00:00ZwithAccept: application/fhir+jsonandPrefer: respond-asyncand a Bearer token. The server validates the token, parses parameters, opens an export job in shared state, returns202 AcceptedwithContent-Location: https://fhir.example.org/export-status/abc. The handler does no actual data extraction in line; it returns within milliseconds.Polling. The client polls
GET /export-status/abcperiodically (honoringRetry-After). While the job runs, the server returns202 Accepted, optionally with anX-Progressheader carrying a free-form message. When the job is finished, the server returns200 OKwith the JSON manifest above. The client now has every URL it needs to fetch the data.Download. The client fetches each
output[].urlin parallel. The server streams each file asapplication/fhir+ndjson, optionallyContent-Encoding: gzip. The number of files, their sizes, and the order are server-chosen.Cleanup. When the client is done - or whenever it decides to abandon the export - it sends
DELETE /export-status/abc. The server returns202 Accepted. The files may now be reclaimed. The status URL begins returning404 Not Found.Notice three things in that diagram. The kick-off handler does no extraction. The polling client may land on a different HFS instance than the one that received the kick-off - the status URL must work regardless. The download path is not necessarily served by HFS at all; if the manifest's
requiresAccessTokenisfalse, the URLs may point directly at the object store.These are not implementation details. They are the architectural premises the rest of this document is responding to.
The Architectural Tensions
Before we get to traits, it is worth naming the tensions that the design has to resolve. Bulk export is not a single piece of code; it is a system, and the system has to hold together under four pressures simultaneously.
Long-running work behind a short-lived HTTP request. The kick-off responds
202in milliseconds. The job behind it may run for minutes, hours, or - for the largest population-level pulls - long enough to outlive the process that started it. The handler and the worker cannot be the same thing.State that outlives a process. Job status, cursors, manifests, output files - all of it has to survive process restarts, deploys, and crashes. There is no "in-memory only" version of an export that is also production-grade. Whatever state we keep must be durable from the first call to
start_export().One server versus many. A small clinic might run HFS as a single process on a single VM. A national exchange will run HFS as a fleet of pods behind a load balancer, scaled to traffic. The kick-off, the status polls, and the file downloads will land on whatever instance the load balancer picks at the moment. Job state cannot live in one instance's memory if any other instance might field the next request.
The download endpoint is a fileserver. Once a job is done, every output URL is a sustained GET. Megabytes per file, gigabytes per export, hundreds of files in the manifest. That is a fundamentally different workload from "look up a Patient by ID and return JSON" - it is hot-path bandwidth, not request/response latency. Bolting it onto the same Tokio executor that fields
Patient.readqueries is workable in a single-instance deployment and a mistake at scale.The design that follows pulls these four tensions apart cleanly, so each one is solved by a single, replaceable abstraction.
Single-Instance vs Multi-Instance: A Tale of Two Deployments
HFS has always tried to scale down as well as it scales up. The persistence layer ships with a zero-config SQLite default; the same trait surface accepts PostgreSQL, MongoDB, Elasticsearch, and S3. Bulk export follows the same philosophy: the same traits serve both the single-VM clinic install and the multi-pod cloud deployment, and the operator decides at startup which concrete implementations to wire in.
Single-instance: zero-config
The simplest possible export deployment is a single HFS process, running on a single VM, writing job state into the SQLite database it already manages, and writing output files to the local filesystem under
${HFS_DATA_DIR}/exports/{job_id}/. The worker that performs the extraction is a Tokio task spawned from the same process; the polling and download handlers serve the local SQLite row and the local file directly.This is fine for clinics, single-tenant deployments, demos, conformance testing, and CI. It works without an external job queue, without object storage, without a network filesystem. The trade-off is that you cannot horizontally scale HFS - the moment a second pod appears behind the load balancer, status polls will start landing on the wrong instance, and the design falls apart.
Multi-instance: shared state, work pool
The horizontally scaled deployment splits responsibilities cleanly:
bulk_export_jobstable. Status polls work from any instance because every instance is looking at the same row.requiresAccessToken: falsecase, the client downloads pre-signed URLs directly from the object store and HFS is not in the path at all.hfs-exporterbinary against the same shared state (an option discussed later). Workers claim jobs out of the shared store using a leasing pattern, so adding or removing workers requires no coordination beyond the shared store itself.The cardinal architectural rule is that the same code path serves both topologies. The
BulkExportStoragetrait is implemented by an embedded SQLite backend for single-instance and a PostgreSQL backend for multi-instance. TheExportOutputStoretrait is implemented by a local-FS backend for single-instance and an S3 backend for multi-instance. The handler does not know which one is wired up. The worker does not know which one is wired up. Only the bootstrap code, reading environment variables, knows.The Recommendation: PostgreSQL for Job State, S3 for Output, In-Process Workers
There is a tension between offering a menu of options and recommending a default. Discussion #45 leaned toward recommendations -
JwksBearerAuthProvideras the default token validator, withIntrospectionAuthProvideras the fallback for opaque tokens. We take the same posture here.For multi-instance job state, the default is PostgreSQL. PostgreSQL is already a supported HFS primary store, so adopting it for export state adds no new operational dependency for the common case.
SELECT ... FOR UPDATE SKIP LOCKEDis the canonical pattern for transactional job queuing in Rust ecosystems (sqlx,tokio-postgres, every Sidekiq-style job library); it handles worker fail-over, lease expiry, and at-least-once delivery without an external broker. Thebulk_export_jobstable is small, write-amplified per heartbeat, and bounded in size by output retention - it does not threaten the resource store's hot path.For multi-instance output storage, the default is S3-compatible object storage. This is the same S3 backend the persistence layer already ships, with the same
AwsS3Clientand the same keyspace conventions. Output keys are scoped under/{tenant}/exports/{job_id}/{resource_type}-{part}.ndjson. The manifest publishes pre-signed URLs with a configurable TTL, so the manifest'srequiresAccessTokenisfalseand the client downloads directly from S3 without HFS in the bandwidth path. For deployments that want to keep token-based access (audit-heavy environments, environments without a CDN), the same files are streamable through HFS's own download handler withrequiresAccessToken: trueinstead - configuration, not code.For execution, the default is an embedded worker pool. Workers run in the same process as the HFS REST API by default. This keeps the operational surface small: one binary, one deployment, one set of logs and metrics. A configurable
HFS_BULK_EXPORT_WORKER_CONCURRENCYlimits how many jobs each pod runs at once, andHFS_BULK_EXPORT_DISABLE_LOCAL_WORKER=truelets operators turn off in-pod workers entirely when they want to dedicate request-serving capacity. The optionalhfs-exporterbinary, discussed later, addresses the cases where worker isolation needs to be physical, not just configurational.Now the vendor-style walkthroughs, in the same shape as the IdP integration section of Discussion #45.
PostgreSQL (recommended default for job state)
How it connects. The same
HFS_DATABASE_URLthat drives the persistence layer. The export subsystem adds two tables -bulk_export_jobsandbulk_export_outputs- alongside the existing resource schema. No new connection pool, no new credentials, no new operational surface.Trade-offs. PostgreSQL is durable, transactional, and well understood.
SELECT ... FOR UPDATE SKIP LOCKEDmakes multi-worker claiming straightforward and safe. The cost is that every status poll is a query against PostgreSQL, which on a hot system means tuning indexes on(tenant_id, status, lease_expiry)and accepting that very-high-poll-rate workloads (more than a few thousand polls per second per tenant) will eventually want a caching layer in front.Configuration sketch.
Redis (alternative for low-latency status polls)
How it connects. A standard Redis or Redis Cluster endpoint. Job records are hashes keyed by job ID; an indexed sorted-set holds pending jobs ordered by enqueue time; claim is
BLMOVEfrompendingtoin-flight-{worker_id}lists with a TTL-backed lease.Trade-offs. Redis makes status polls trivially fast - a single
HGETALLis under a millisecond - and the claim semantics are clean. The cost is durability: a Redis crash without AOF persistence can lose in-flight job state. For exports, the worst-case impact is a job that has to be restarted; cursors live in PostgreSQL via the sameBulkExportStoragetrait, so re-running is mostly idempotent, but operators who run Redis as a cache rather than a primary store should think twice. Best fit: deployments that already operate Redis as a hot path and want polling latency under a millisecond.Configuration sketch.
DynamoDB / Cosmos DB / Spanner (cloud-managed equivalents)
How they connect. Each cloud's identity model. DynamoDB via the AWS SDK; Cosmos DB via the Azure SDK; Spanner via Google Cloud credentials. Each is a
BulkExportStorageimplementation that mirrors the PostgreSQL pattern but uses conditional writes (DynamoDB:ConditionExpression, Cosmos: ETag preconditions, Spanner: read-modify-write transactions) in place ofSKIP LOCKED.Trade-offs. Managed durability and global replication, at the cost of additional integration code per provider and per-call billing. The same caveats as the IdP discussion in #45 apply: every cloud has its own claim-name and capability quirks; the abstraction has to absorb them.
These are not first-tier targets for the initial implementation, but the trait design must not preclude them. Anyone running HFS purely on a single cloud will eventually want them.
Kafka / NATS JetStream (workers physically separate from request handlers)
How they connect. Kafka topics or JetStream streams act as the work queue; HFS publishes a job-created event on kick-off, and a separate
hfs-exporterbinary consumes the topic. Job state still lives in PostgreSQL (or wherever theBulkExportStorageimpl points), so status polls and downloads do not touch the broker.Trade-offs. This is the "we run our exporters on a different node pool because they're bandwidth-heavy" case. Adds a broker as an operational dependency, but lets you scale request handlers and exporters independently, and gives you explicit at-least-once delivery semantics with offsets. Best fit: fleets large enough that the export workload visibly distorts request-serving capacity.
S3-compatible object storage (recommended default for output files)
How it connects. The same
AwsS3Clientthe persistence layer's S3 backend uses, configured via the standard AWS credential chain. Output files are uploaded as multipart objects under/{tenant}/exports/{job_id}/{resource_type}-{part}.ndjson. The manifest publishes pre-signedGETURLs with a TTL configured byHFS_BULK_EXPORT_FILE_URL_TTL.Trade-offs. Object storage is the right tool for the job - massively parallel reads, transparent CDN integration, region-redundant durability, lifecycle policies for automatic expiry. The only meaningful cost is that pre-signed URLs reveal that something exists at this URL until this expiry, which some audit regimes treat as out of band. Those deployments switch
HFS_BULK_EXPORT_REQUIRES_ACCESS_TOKEN=true, the manifest reportsrequiresAccessToken: true, and downloads flow through HFS's own handler instead.Cloudflare R2 / Google Cloud Storage / MinIO (S3-compatible drop-ins)
R2, GCS (via interop), and MinIO all speak the S3 API. The same
AwsS3Clientworks against each with no code change, onlyendpoint_urlandforce_path_styleconfiguration adjustments. We will document adocker composeexample with MinIO as part of the development environment so contributors can exercise the multi-instance path without an AWS account.Local filesystem (single-instance only)
How it connects. Files are written to
${HFS_DATA_DIR}/exports/{tenant_id}/{job_id}/{resource_type}-{part}.ndjson. The download handler serves them viatokio::fs::Fileand Axum's streaming body.Trade-offs. No external dependencies; perfect for development and single-VM deployments. Not safe in a multi-instance topology because the writing instance and the reading instance may differ. A shared NFS mount makes this technically work across instances, but it is brittle (lock semantics, cache coherency, fsync surprises) and we do not recommend it.
Designing the Rust Traits
The persistence crate already carries most of the building blocks. The export module at
crates/persistence/src/core/bulk_export.rsdefines the types and traits below; the S3 backend implements them today, and the embedded SQLite and PostgreSQL backends will follow. We present the existing surface first - so readers know what already exists - and then propose the additions that this discussion is centrally about.The Existing Surface: Types
The vocabulary of an export job. These are stable; nothing in this document proposes changing them.
ExportProgressand its per-type companionTypeExportProgressround out the model.TypeExportProgresscarries the cursor state that lets a worker resume mid-job after a crash - the same cursor type thatExportDataProvider::fetch_export_batchaccepts and returns.ExportManifestandExportOutputFilemodel the terminal manifest that the status endpoint serves at200 OK.NdjsonBatchis what data providers produce - one logical batch of NDJSON lines plus a cursor and a "is this the last batch" flag.The Existing Surface: Job-State Trait
BulkExportStorageis the contract that any backend providing job state implements. Single-instance backends (embedded SQLite) and multi-instance backends (PostgreSQL, Redis) both satisfy this trait. The handler does not care which is wired up.The Existing Surface: Data-Provider Trait Hierarchy
ExportDataProvideris what the worker calls when it needs more resources to write. It is notBulkExportStorage- one provides job lifecycle, the other provides data. In practice the same backend object often implements both (the persistence layer already does), but conceptually they are separate concerns that could live in separate processes.The trait hierarchy reflects the spec: every server that can do a Patient export can also do a System export; every server that can do a Group export can also do a Patient export. The compiler enforces the capability ladder.
The Proposal: Output Storage as a First-Class Trait
The traits above answer "what data do we export" and "how is the job tracked". They do not answer "where do the bytes go". Today, each backend that implements
BulkExportStoragealso implicitly decides where output files live - the S3 backend writes them to S3, a future SQLite backend would write them locally. This works, but it conflates two decisions that operators reasonably want to make independently. A site running PostgreSQL for job state and S3 for output is a perfectly normal configuration. A site running PostgreSQL for both is also reasonable. The current shape does not support that cleanly.We propose a separate
ExportOutputStoretrait:Two implementations cover the common cases.
LocalFsOutputStorewrites under${HFS_DATA_DIR}/exports/, hands back HFS-served URLs, and setsrequires_access_token: true.S3OutputStorewrites to a bucket configured via the existingS3BackendConfig, hands back pre-signed URLs, and setsrequires_access_token: false(ortrueif the operator wants to keep downloads on the HFS data path).The Proposal: Workers, Leases, and the Claim Strategy
Today, the S3 backend's
start_exportdoes the work synchronously inside the kick-off handler. That works for the single-instance case but breaks two design goals: it pins the handler thread for the duration of the job, and it gives no way to scale workers independently of request handlers. The handler must return immediately; the work happens elsewhere.A worker is the runtime that performs an export. Workers may run in-pod (the default) or in a separate
hfs-exporterbinary. Either way, they share state throughBulkExportStorageand they coordinate through a leasing protocol.The choice of claim mechanism is itself pluggable, so the same
ExportWorkerruntime works against any job-state backend:Three implementations are in scope for the initial work:
PostgresSkipLocked(default for multi-instance),RedisListMove(alternative for low-poll-latency deployments), andInMemoryMutex(used by the embedded single-instance backend and by tests).Two design notes that come up in review:
Why a lease with expiry rather than an explicit ack/nack queue? Because the work is long-lived and idempotent (cursors live in
TypeExportProgress). A worker that dies mid-job leaves its lease to expire; another worker picks the job up from the last persisted cursor. This is simpler to operate than a queue with explicit redelivery, and it matches howtokio-postgresjob-queue libraries already work.Why a fencing token? Because the lease-expiry pattern allows two workers to briefly believe they hold the same job if the original worker hung rather than crashed. The fencing token, written into every output-file key and checked on the output store's
finalize_part, prevents the zombie worker from corrupting output the live worker is producing. Inspired directly by the fencing tokens pattern Martin Kleppmann wrote about; nothing novel.The Proposal: File-Download Authorization
The download endpoint is its own small authorization problem. The manifest can be served two ways -
requiresAccessToken: true, meaning download URLs point at HFS and require the kick-off's Bearer token; orrequiresAccessToken: false, meaning download URLs point at the object store and are pre-signed. The download handler in HFS needs to handle the first case; the second case bypasses HFS entirely.The default implementation is
BearerScopeAuth. It revalidates the token against the sameAuthProviderdiscussed in Discussion #45, checks that thesystem/{ResourceType}.rsscope covers the file's resource type, and lets the handler stream the file. The pre-signed URL case never runs through this trait - by the time a client is downloading a pre-signed URL, the object store is doing the auth check via the URL's signature.The REST Layer: How the Endpoints Wire Up
Four handlers, all in the established HFS style: generic over the storage trait, taking a
TenantExtractor, returningRestResult<Response>. We sketch the kick-off here; the other three follow the same pattern.Three observations on this handler that are easy to miss:
It does not spawn a Tokio task to run the export. The
start_exportcall writes the job row and returns. A worker picks the job up viaclaim_nextout of band. The handler returns within milliseconds even for jobs that will take hours.It does not know whether the deployment is single-instance or multi-instance.
state.storage()is whicheverBulkExportStoragewas wired in at startup; the handler is identical either way.It runs inside the auth middleware described in Discussion #45. By the time this function runs,
ctxis a fully validatedRequestContextcontaining aPrincipalwith aScopeSet. The handler does not re-validate the token; it only callsauthorize_kickoffon whateverAuthorizationPolicyis configured, which evaluates SMART system scopes (and any composed deployment-specific policies) against the requested export.The status handler is symmetric:
Content-LocationURLs are constructed fromHFS_BASE_URLplus the tenant prefix (whenHFS_TENANT_ROUTING_MODE=url_path) plus/export-status/{job_id}. They are absolute. They survive load balancer changes because every HFS instance constructs the same URL for the same job, and every instance can answer the poll against the sharedBulkExportStorage.Error semantics. Per-resource-type failures during the run are not catastrophic - they accumulate as
OperationOutcomeresources inerror[]NDJSON files attached to the manifest, and the job still terminatesComplete. Only conditions that prevent the export from producing any valid output (authorization failure mid-stream, total backend outage, output-store failure on every write) transition the job toExportStatus::Error. This matches the IG's expectation that bulk jobs prefer partial success over hard failure.Group Export: The Hard Part
Group/[id]/$exportis where the spec's edges become apparent. The export returns "every resource in the patient compartment of every patient who is a member of this group" - which is the cross product of three things HFS has to compute on the fly.First, who are the members? Groups can list members directly, list nested Groups whose members must be flattened, or (in the forthcoming Bulk Cohort profile) carry
member-filtermodifier extensions whose values are FHIR search expressions to evaluate against the live data.GroupExportProvider::get_group_membershandles the direct case;resolve_group_patient_idshandles the rest.Second, what is the patient compartment for this FHIR version? Compartments are defined per FHIR version by
CompartmentDefinitionresources; the mapping from(version, resource_type)to(search_param_names)is generated alongside the FHIR models. HFS already has this lookup atcrates/rest/src/handlers/compartment.rs::get_compartment_params_for_version. The bulk export worker reuses it.Third, how do we enumerate efficiently? For each requested resource type, the worker calls
fetch_patient_compartment_batchwith the resolved patient ID list and a cursor. The implementation chooses whether to issue per-patient queries, range queries, or a single chunked query depending on what the underlying store is good at; the trait does not prescribe.The IG's behavior on
_sinceplus group membership has a subtle wrinkle that operators should be aware of: if a patient was added to the group after_since, the server MAY return that patient's resources from before_since(because they were not part of the group at that time, but are now). The current draft says "behavior SHOULD be documented" - we will document our choice (we plan to include them by default, matching the prevalent reading) and let operators override viaHFS_BULK_EXPORT_SINCE_NEWLY_ADDED=excludewhen their use case demands the alternative.The Bulk Cohort
member-filterprofile - where a client can POST a Group whose membership is defined by FHIR search criteria evaluated server-side - is intentionally out of scope for the first cut. It is its own design problem (asynchronous Group construction, dynamic membership, refresh semantics) and deserves its own discussion document once the export plumbing is in production.Authorization
Bulk export's authorization story is short because Discussion #45 has already done the work.
By the time an export handler runs, the auth middleware has produced a
RequestContextwith a validatedPrincipaland aScopeSet. The export kick-off handler asks the sameAuthorizationPolicytrait whether the principal's scopes cover the requested export. The scopes that matter are the standard SMART Backend Services system scopes:system/*.rssystem/Patient.rssystem/Observation.rssystem/[Type].readThe composability described in #45 carries through here. A deployment that wants additional restrictions on bulk operations - say, a
BulkRateLimitPolicythat throttles concurrent jobs per client, or aBulkTenantQuotaPolicythat caps total exported volume per tenant per day - implementsAuthorizationPolicyand composes it viaCompositeAuthorizationPolicy. The export handlers do not know these policies exist; they only know thatauthorize_kickoffreturnedPermit.What this means in practice: there is no new auth surface for bulk export. The same
JwksBearerAuthProvideryou configured for the rest of HFS validates the kick-off token. The same scope syntax governs what can be exported. The same audit trail records every job.Configuration: What Operators Will Touch
HFS_BULK_EXPORT_ENABLEDtruefalse, the operation endpoints return501 Not Implemented.HFS_BULK_EXPORT_BACKENDembeddedembedded(SQLite + local FS),postgres-s3,redis-s3.HFS_BULK_EXPORT_DATABASE_URLHFS_DATABASE_URL)HFS_BULK_EXPORT_OUTPUT_BACKENDlocal-fslocal-fsors3.HFS_BULK_EXPORT_OUTPUT_DIR${HFS_DATA_DIR}/exportsHFS_BULK_EXPORT_S3_BUCKETOUTPUT_BACKEND=s3.HFS_BULK_EXPORT_REQUIRES_ACCESS_TOKENautoauto(pre-signed when supported),true(always token),false(always pre-signed).HFS_BULK_EXPORT_FILE_URL_TTL3600HFS_BULK_EXPORT_OUTPUT_TTL86400HFS_BULK_EXPORT_WORKER_CONCURRENCY2HFS_BULK_EXPORT_DISABLE_LOCAL_WORKERfalsetrue, this pod does not run workers (use with separatehfs-exporter).HFS_BULK_EXPORT_MAX_CONCURRENT_PER_TENANT4HFS_BULK_EXPORT_BATCH_SIZE1000fetch_export_batchcall.HFS_BULK_EXPORT_LEASE_DURATION60HFS_BULK_EXPORT_HEARTBEAT_INTERVAL20HFS_BULK_EXPORT_SINCE_NEWLY_ADDEDincludeincludeorexcluderesources from before_sincefor patients added after_since.The single-instance default -
HFS_BULK_EXPORT_BACKEND=embeddedwithHFS_BULK_EXPORT_OUTPUT_BACKEND=local-fs- requires zero additional configuration on top of the standard HFS environment. A deployment grows into the multi-instance path by changing two variables and pointing at a Postgres and an S3 bucket; no code changes, no different binary.Conformance Testing
The Inferno Bulk Data Test Kit is the canonical conformance harness for FHIR bulk data servers. It exercises every kick-off variant, the polling pattern, the manifest schema, the NDJSON output format, the cancellation flow, and SMART Backend Services authorization end to end. After the initial implementation lands, we will:
cargo xtask inferno-bulk-datatarget that runs the test kit headlessly against this configuration..github/workflows/inferno.ymlalongside the existing test kits.crates/hfs/README.mdnext to the other test-kit badges.This is intentionally separate from unit and integration tests in the workspace. Inferno tests are slow, network-bound, and authoritative - they belong in CI as a nightly job, not on every PR.
What's Not in Scope (Yet)
A handful of things are deliberately deferred. None are blockers; each is a follow-up.
$bulk-submit(the inverse direction - large NDJSON payloads into the server). The Argonaut Project's current draft is at https://hackmd.io/@argonaut/rJoqHZrPle. It will get its own discussion document once the draft stabilizes; the shared-state architecture proposed here generalizes naturally to ingestion.member-filterprofile for dynamic Group construction. This is its own design problem (asynchronous Group creation, dynamic membership, refresh semantics) and deserves its own discussion.$importfrom earlier IG drafts. Superseded by$bulk-submit; we will not implement the legacy shape.Prefer: separate-export-status(the variant where status polling returns200 OKwith anX-Export-Statusheader instead of202 Accepted). Marked as a follow-up; it is a low-effort addition once the core flow is in place.organizeOutputBy(reorganized output with Parameters header blocks per group). Wait for broader IG adoption before committing to it.includeAssociatedData=LatestProvenanceResourcesand similar Provenance hints. Implement once the audit subsystem's Provenance support lands.Proposed Next Steps
The traits sketched above are a starting point. To move toward implementation:
ExportOutputStoreandExportClaimStrategytohelios-persistence. Two new traits; refactor the existing S3 bulk-export code to satisfy them rather than implementing everything inside the S3 backend'sBulkExportStorage.PostgresSkipLockedand the PostgreSQLBulkExportStorage. With the S3ExportOutputStore, this is the multi-instance default. Cover with testcontainers integration tests against real Postgres + MinIO.helios-rest. Four handlers: kick-off (with sub-routes for system, patient, group), status, cancel/delete, file-download. PlumbContent-Location,X-Progress,Retry-After,Expires, and the manifest content type correctly.AwsS3Clientalready supports it via the AWS SDK.docker composeconfiguration with HFS + PostgreSQL + MinIO + Keycloak, suitable for running the Inferno Bulk Data Test Kit locally and in CI..github/workflows/inferno.ymlas a nightly conformance job, and publish the badge.HFS_BULK_EXPORT_*envvars inCLAUDE.mdandcrates/hfs/README.md, including the single-instance vs multi-instance configuration recipes.record_export_eventhelper already incrates/persistence/src/core/bulk_export.rs::audit, plumbed through the kick-off, completion, cancellation, and download handlers.$bulk-submitdiscussion once the Argonaut draft is stable, building on the shared-state architecture established here.Closing Thoughts
Bulk export is the API that turns a FHIR server from a transaction processor into a data platform. Population health teams, research data lakes, payer-provider exchanges, AI training pipelines - none of them are reading one Patient at a time. They are reading entire compartments, entire cohorts, entire systems, and they want to do it asynchronously, resumably, and at a rate that does not require an arrangement with the FHIR server's on-call.
The architecture proposed here is built around two convictions. First, that the same trait surface should serve a single VM with SQLite and a horizontally scaled fleet with PostgreSQL and S3 - the operator chooses, the code does not change. Second, that the long-running, bandwidth-heavy parts of an export should be cleanly separable from the request-serving HFS process, so that operators can scale them independently without rewriting handlers.
The Rust trait system makes both convictions enforceable. The compiler guarantees that every export handler receives a validated
RequestContextand aTenantContext. ThatBulkExportStorage,ExportDataProvider, andExportOutputStoreare independently replaceable. That the worker runtime is identical whether it is co-located with the REST API or running standalone. These guarantees hold regardless of how complex the deployment becomes, and they hold across the inevitable migrations from "we started on SQLite and outgrew it" to "we now run on Postgres + S3 + a separate worker tier".After the implementation lands, the Inferno Bulk Data Test Kit becomes the daily check on whether HFS is a conformant bulk data server. The kit covers every kick-off variant, every flavor of the polling state machine, every required manifest field, the NDJSON contract, and SMART Backend Services authorization end to end. Treating Inferno conformance as a non-negotiable in CI is what turns "we shipped bulk export" into "we shipped a bulk export implementation interoperable with the rest of the ecosystem".
Thank you for reading. I look forward to the discussion.
All reactions