Skip to content

ci(bench): benchmark all six storage backends (#1475) - #1544

Merged
smunini merged 4 commits into
ci/bench-host-hardeningfrom
ci/1475-benchmark-backend-matrix
Sep 28, 2026
Merged

smunini merged 4 commits into
ci/bench-host-hardeningfrom
ci/1475-benchmark-backend-matrix

Conversation

@dougc95

@dougc95 dougc95 commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Closes #1475. Stacked on #1543, which fixes shared-host cleanup problems that already exist on main. This PR's base is that branch, so the diff below is #1475 only. GitHub retargets it to main once #1543 merges.

What it does

fhir-benchmark.yml can now benchmark every storage backend, each with and without Elasticsearch:

Leg Port Stateful containers (remote Docker host)
sqlite 8081 none
postgres 8082 postgres:18
mongodb 8083 mongo:7.0, single-member replica set (FHIR transaction bundles require one)
sqlite-elasticsearch 8084 elasticsearch:8.15.0
postgres-elasticsearch 8085 postgres:18 + elasticsearch:8.15.0
mongodb-elasticsearch 8086 mongo:7.0 + elasticsearch:8.15.0

Bare s3 is left out because it has no search (s3/storage.rs), so import's conditional references and the search suite can't run on it. s3-elasticsearch is a proposed follow-up.

Behaviour changes for existing dispatchers

  • The backend input's default is now core, which means sqlite + postgres, the same legs as today. all now means all 6 legs. The new elasticsearch choice runs the 3 composite legs, and any single backend can be chosen on its own.
  • Legs now run one at a time by default (max_parallel, default 1, maximum 2). Every leg shares one 12 GB / 4-CPU Docker host, and past runs saw neighbour load move throughput 2–17×. A default run's wall time therefore roughly doubles.
  • cancel-in-progress can now cancel a run of up to 6 legs.
  • benchcmp.py lives outside this repo. It may need to learn the new backend names and the new runner-info.txt keys.

New inputs

These are validated in setup (every bad value is reported) before the ~13-minute build starts.

Input Default Notes
max_parallel 1 1..2
es_heap 1g 256m..4g
es_sync_mode asynchronous asynchronous is HFS's default: acknowledged on primary commit, then forwarded to ES by one worker. synchronous batches with _bulk and uses wait_for.
mongo_wt_cache_gb 2 0.25..6. The same budget as Postgres shared_buffers. mongod's own default (~5 GB) risks a host OOM.
hfs_mongo_max_connections 32 HFS's default of 10 starves crud's 300 VUs

Measurement safeguards

The new legs could otherwise produce numbers that look comparable when they aren't.

  • ES drain gate, before the search suite: a conditional DELETE /Patient?identifier=<no match> goes through ensure_writes_visible → SyncManager::barrier(), so it returns only after the async sync worker has processed every earlier write. The gate then waits for the ES counters to settle, and finally compares live resource counts in the primary with top-level ES documents. The result goes to es-drain.txt (drained / incomplete / timeout).
  • Import completeness: if fewer than 1000 bundles were imported, the leg's crud and search results are flagged as not comparable.
  • Result-size cross-check: after the search suite, 12 _summary=count queries are run, alongside what crud and prewarm left behind (crud-residue.txt). This catches backends returning different result sets.
  • Capacity gate: each leg waits up to 10 minutes for enough MemAvailable on the Docker host (primary + ES memory + a 2 GB margin), then fails rather than risk an OOM. Input combinations that can never fit fail immediately.
  • Summary: each leg gets a tuning table, backend stats, and a "How to read this leg" note (on async ES legs, import and crud throughput is the sync worker's rate). Mongo transaction-abort counts are recorded, and a backend container that dies fails the leg.

Shared-host hygiene (builds on #1543)

  • Every container carries the labels hfs-bench=1, hfs-bench-run, hfs-bench-leg and hfs-ci=true, plus --log-driver local. Every named volume is labelled.
  • Every container has a Stop step and is covered by the leak check.
  • Helper containers are removed even when timeout kills the Docker CLI.

Layout

The new bash and Python live in .github/scripts/fhir-bench/, following the pattern of .github/scripts/obs-ab/:

  • resolve-matrix.sh, capacity-gate.sh, start-mongodb.sh, start-elasticsearch.sh, diagnose.sh;
  • suite-lib.sh, sourced by "Run benchmark suites";
  • crud_residue.py and summary_backends.py.

The workflow keeps the orchestration. Each script's header lists the env vars it reads and the files or env it writes.

Verification

  • Static checks:
    • actionlint 1.7.12: clean.
    • shellcheck 0.11.0 on every script: clean.
    • bash -n and Python ast checks: clean.
    • cargo check -p helios-hfs --no-default-features --features "R4,sqlite,postgres,mongodb,elasticsearch" passes. CI doesn't build this feature set anywhere else.
  • Setup script: run locally for every backend value and for edge-case max_parallel, es_heap, mongo_wt_cache_gb and hfs_mongo_max_connections values.
  • Review: three adversarial review rounds (Actions/bash, shared-host safety, benchmark validity against the HFS code), then an equivalence review confirming that extracting the scripts changed no behaviour.
  • Live runs, results appended as they finish:
    • B1 — 36410157709: -f tests=prewarm -f max_parallel=2, the core legs run together (names, capacity gate, leak check). pending
    • B2 — async sqlite-elasticsearch with tests=prewarm,import,search. This decides the default: if fewer than 1000 bundles import, es_sync_mode defaults to synchronous in this PR. pending
    • Later:
      • all + prewarm,search;
      • mongodb-elasticsearch synchronous with import;
      • invalid inputs;
      • a full all run.

🤖 Generated with Claude Code

https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ

dougc95 and others added 3 commits September 28, 2026 06:18
Extends fhir-benchmark.yml from sqlite + postgres to every storage
backend with and without Elasticsearch: sqlite, sqlite-elasticsearch,
postgres, postgres-elasticsearch, mongodb and mongodb-elasticsearch.
Bare s3 is left out because it has no search, so the import and
search suites cannot run on it; s3-elasticsearch is a follow-up.

Dispatch:
- backend: core (default; sqlite + postgres), all (6 legs),
  elasticsearch (the 3 composites), or any single backend.
- max_parallel (default 1, max 2): legs share one 12 GB / 4-CPU
  Docker host, so running them together skews every leg's numbers.
- es_heap, es_sync_mode (asynchronous = HFS default | synchronous),
  mongo_wt_cache_gb, hfs_mongo_max_connections. The setup job
  validates them all before the build starts.

Per leg:
- MongoDB runs as a single-member replica set (transaction bundles
  need one), with a capped WiredTiger cache and a 900 s transaction
  lifetime to match HFS_REQUEST_TIMEOUT.
- Elasticsearch runs single-node. Yellow health is expected, because
  HFS creates every index with one replica.
- A capacity gate waits up to 10 minutes for enough free memory on
  the Docker host and fails the leg rather than risk an OOM.
- Containers and volumes carry the hfs-bench, hfs-ci and leg labels,
  so the reaper, docker-host-gc and the leak check all cover them.

Measurement:
- On ES legs, a drain gate runs before the search suite: a
  conditional DELETE that matches nothing acts as a barrier on the
  composite's sync queue, then ES counters must settle, then live
  primary and ES resource counts are compared.
- Import completeness, per-type crud leftovers and a _summary=count
  cross-check are recorded, so legs that did not load the same data
  are flagged rather than silently compared.
- Backend stats, a "how to read this leg" note and the tuning knobs
  go into the step summary and runner-info.txt.

The new bash and Python live in .github/scripts/fhir-bench/, following
.github/scripts/obs-ab, so the workflow keeps its orchestration and
the scripts can be linted directly.

Closes #1475

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ
Run 36410157709 showed the gate measuring the wrong machine:
`docker info` reported MemTotal 12000 MB while a plain container's
/proc/meminfo showed MemAvailable 61324 MB, so the gate would pass
almost every time. The Docker daemon evidently runs inside a 12 GB
limit that containers' /proc/meminfo does not reflect (likely an
lxcfs-virtualised LXC, where even a bind-mounted meminfo would report
free memory relative to the reading container's own cgroup).

host-mem.sh now derives available memory from one source it can
trust: docker info MemTotal minus the usage `docker stats` reports for
every running container, minus 1024 MB for dockerd and non-container
processes. If `docker stats` fails the reading is `unknown` rather
than "nothing running", so the gate keeps polling and fails at its
10-minute budget instead of passing on a guess.

The capacity gate, probe_host (host_mem_avail_mb + host_mem_source in
host-contention.txt), runner-info.txt (capacity_mem_source), the step
summary and diagnose.sh all use it. probe_host now runs before
SUITE_START, so suite wall time no longer includes the probe itself.

Refs #1475

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ
Run 36412022609 (sqlite-elasticsearch, asynchronous, prewarm+import+
search) showed that async mode cannot finish the import corpus inside
k6's 60-minute cap: 313 of 1000 bundles succeeded, 14 hit the 900 s
client timeout, and bundle latency p95 was 489 s. The drain gate then
reported status=incomplete: SQLite held 562,649 live resources and
Elasticsearch 367,985, with composite_secondary_sync_needs_reindex=0.

As agreed for this case, es_sync_mode now defaults to synchronous
(bundles batched via _bulk, refresh=wait_for). asynchronous, HFS's own
out-of-the-box mode, stays available as an input, and its
"How to read this leg" note now says import is expected to hit the cap.

Refs #1475

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ
Dispatching cffc076 failed before any job ran: "failed to parse
workflow: (Line: 1145, Col: 14): Exceeded max expression length
21000". A run: block that contains any ${{ }} is evaluated as one
expression, and GitHub caps expressions at 21,000 characters. The
"Run benchmark suites" block had grown to ~26,300 characters with 31
substitutions.

Every matrix/inputs/github/needs value that block uses now comes in
through the step's env: (BACKEND, BENCH_PORT, BENCH_RUN_ID,
BENCH_IN_*, BENCH_SHA, BENCH_REF_NAME, BENCH_MAX_PARALLEL,
BENCH_LEG_TIMEOUT_MIN), with the same `||` fallbacks done in bash, so
the block is a plain string with no length cap. A comment above the
env: block records the rule. No behaviour change.

Refs #1475

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ
@smunini
smunini merged commit 15fdaca into ci/bench-host-hardening Sep 28, 2026
2 of 8 checks passed
@smunini
smunini deleted the ci/1475-benchmark-backend-matrix branch September 28, 2026 13:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants