ci(bench): benchmark all six storage backends (#1475) - #1544
Merged
smunini merged 4 commits intoSep 28, 2026
Merged
Conversation
Extends fhir-benchmark.yml from sqlite + postgres to every storage backend with and without Elasticsearch: sqlite, sqlite-elasticsearch, postgres, postgres-elasticsearch, mongodb and mongodb-elasticsearch. Bare s3 is left out because it has no search, so the import and search suites cannot run on it; s3-elasticsearch is a follow-up. Dispatch: - backend: core (default; sqlite + postgres), all (6 legs), elasticsearch (the 3 composites), or any single backend. - max_parallel (default 1, max 2): legs share one 12 GB / 4-CPU Docker host, so running them together skews every leg's numbers. - es_heap, es_sync_mode (asynchronous = HFS default | synchronous), mongo_wt_cache_gb, hfs_mongo_max_connections. The setup job validates them all before the build starts. Per leg: - MongoDB runs as a single-member replica set (transaction bundles need one), with a capped WiredTiger cache and a 900 s transaction lifetime to match HFS_REQUEST_TIMEOUT. - Elasticsearch runs single-node. Yellow health is expected, because HFS creates every index with one replica. - A capacity gate waits up to 10 minutes for enough free memory on the Docker host and fails the leg rather than risk an OOM. - Containers and volumes carry the hfs-bench, hfs-ci and leg labels, so the reaper, docker-host-gc and the leak check all cover them. Measurement: - On ES legs, a drain gate runs before the search suite: a conditional DELETE that matches nothing acts as a barrier on the composite's sync queue, then ES counters must settle, then live primary and ES resource counts are compared. - Import completeness, per-type crud leftovers and a _summary=count cross-check are recorded, so legs that did not load the same data are flagged rather than silently compared. - Backend stats, a "how to read this leg" note and the tuning knobs go into the step summary and runner-info.txt. The new bash and Python live in .github/scripts/fhir-bench/, following .github/scripts/obs-ab, so the workflow keeps its orchestration and the scripts can be linted directly. Closes #1475 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ
Run 36410157709 showed the gate measuring the wrong machine: `docker info` reported MemTotal 12000 MB while a plain container's /proc/meminfo showed MemAvailable 61324 MB, so the gate would pass almost every time. The Docker daemon evidently runs inside a 12 GB limit that containers' /proc/meminfo does not reflect (likely an lxcfs-virtualised LXC, where even a bind-mounted meminfo would report free memory relative to the reading container's own cgroup). host-mem.sh now derives available memory from one source it can trust: docker info MemTotal minus the usage `docker stats` reports for every running container, minus 1024 MB for dockerd and non-container processes. If `docker stats` fails the reading is `unknown` rather than "nothing running", so the gate keeps polling and fails at its 10-minute budget instead of passing on a guess. The capacity gate, probe_host (host_mem_avail_mb + host_mem_source in host-contention.txt), runner-info.txt (capacity_mem_source), the step summary and diagnose.sh all use it. probe_host now runs before SUITE_START, so suite wall time no longer includes the probe itself. Refs #1475 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ
Run 36412022609 (sqlite-elasticsearch, asynchronous, prewarm+import+ search) showed that async mode cannot finish the import corpus inside k6's 60-minute cap: 313 of 1000 bundles succeeded, 14 hit the 900 s client timeout, and bundle latency p95 was 489 s. The drain gate then reported status=incomplete: SQLite held 562,649 live resources and Elasticsearch 367,985, with composite_secondary_sync_needs_reindex=0. As agreed for this case, es_sync_mode now defaults to synchronous (bundles batched via _bulk, refresh=wait_for). asynchronous, HFS's own out-of-the-box mode, stays available as an input, and its "How to read this leg" note now says import is expected to hit the cap. Refs #1475 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ
Dispatching cffc076 failed before any job ran: "failed to parse workflow: (Line: 1145, Col: 14): Exceeded max expression length 21000". A run: block that contains any ${{ }} is evaluated as one expression, and GitHub caps expressions at 21,000 characters. The "Run benchmark suites" block had grown to ~26,300 characters with 31 substitutions. Every matrix/inputs/github/needs value that block uses now comes in through the step's env: (BACKEND, BENCH_PORT, BENCH_RUN_ID, BENCH_IN_*, BENCH_SHA, BENCH_REF_NAME, BENCH_MAX_PARALLEL, BENCH_LEG_TIMEOUT_MIN), with the same `||` fallbacks done in bash, so the block is a plain string with no length cap. A comment above the env: block records the rule. No behaviour change. Refs #1475 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1475. Stacked on #1543, which fixes shared-host cleanup problems that already exist on
main. This PR's base is that branch, so the diff below is #1475 only. GitHub retargets it tomainonce #1543 merges.What it does
fhir-benchmark.ymlcan now benchmark every storage backend, each with and without Elasticsearch:sqlitepostgresmongodbsqlite-elasticsearchpostgres-elasticsearchmongodb-elasticsearchBare
s3is left out because it has no search (s3/storage.rs), so import's conditional references and the search suite can't run on it.s3-elasticsearchis a proposed follow-up.Behaviour changes for existing dispatchers
backendinput's default is nowcore, which means sqlite + postgres, the same legs as today.allnow means all 6 legs. The newelasticsearchchoice runs the 3 composite legs, and any single backend can be chosen on its own.max_parallel, default 1, maximum 2). Every leg shares one 12 GB / 4-CPU Docker host, and past runs saw neighbour load move throughput 2–17×. A default run's wall time therefore roughly doubles.cancel-in-progresscan now cancel a run of up to 6 legs.benchcmp.pylives outside this repo. It may need to learn the new backend names and the newrunner-info.txtkeys.New inputs
These are validated in
setup(every bad value is reported) before the ~13-minute build starts.max_paralleles_heap1ges_sync_modeasynchronousasynchronousis HFS's default: acknowledged on primary commit, then forwarded to ES by one worker.synchronousbatches with_bulkand useswait_for.mongo_wt_cache_gbshared_buffers. mongod's own default (~5 GB) risks a host OOM.hfs_mongo_max_connectionsMeasurement safeguards
The new legs could otherwise produce numbers that look comparable when they aren't.
DELETE /Patient?identifier=<no match>goes throughensure_writes_visible→SyncManager::barrier(), so it returns only after the async sync worker has processed every earlier write. The gate then waits for the ES counters to settle, and finally compares live resource counts in the primary with top-level ES documents. The result goes toes-drain.txt(drained/incomplete/timeout)._summary=countqueries are run, alongside what crud and prewarm left behind (crud-residue.txt). This catches backends returning different result sets.MemAvailableon the Docker host (primary + ES memory + a 2 GB margin), then fails rather than risk an OOM. Input combinations that can never fit fail immediately.Shared-host hygiene (builds on #1543)
hfs-bench=1,hfs-bench-run,hfs-bench-legandhfs-ci=true, plus--log-driver local. Every named volume is labelled.timeoutkills the Docker CLI.Layout
The new bash and Python live in
.github/scripts/fhir-bench/, following the pattern of.github/scripts/obs-ab/:resolve-matrix.sh,capacity-gate.sh,start-mongodb.sh,start-elasticsearch.sh,diagnose.sh;suite-lib.sh, sourced by "Run benchmark suites";crud_residue.pyandsummary_backends.py.The workflow keeps the orchestration. Each script's header lists the env vars it reads and the files or env it writes.
Verification
bash -nand Pythonastchecks: clean.cargo check -p helios-hfs --no-default-features --features "R4,sqlite,postgres,mongodb,elasticsearch"passes. CI doesn't build this feature set anywhere else.backendvalue and for edge-casemax_parallel,es_heap,mongo_wt_cache_gbandhfs_mongo_max_connectionsvalues.-f tests=prewarm -f max_parallel=2, the core legs run together (names, capacity gate, leak check). pendingsqlite-elasticsearchwithtests=prewarm,import,search. This decides the default: if fewer than 1000 bundles import,es_sync_modedefaults tosynchronousin this PR. pendingall+prewarm,search;mongodb-elasticsearchsynchronous with import;allrun.🤖 Generated with Claude Code
https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ