Skip to content

ci(bench): stream the per-shape search attribution instead of slurping it - #1547

Open
dougc95 wants to merge 1 commit into
ci/bench-host-hardeningfrom
ci/bench-attribution-stream
Open

dougc95 wants to merge 1 commit into
ci/bench-host-hardeningfrom
ci/bench-attribution-stream

Conversation

@dougc95

@dougc95 dougc95 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Follow-up to #1544, which is now merged into ci/bench-host-hardening (#1543). This PR's base is that branch.

Problem

"Attribute search latency per query shape" ran jq -rs over bench-results/<backend>/search-points.json. That file is k6's --out json stream, one line for every metric of every search request. -s loads the whole stream into a single in-memory array, so memory grows with the file size, not with the ~20 query shapes being summarised.

This step has died from memory exhaustion three times:

  • run 30550776427;
  • run 33515369645;
  • run 36423973848, sqlite leg on github-agent3. The log shows 109920 Killed jq -rs after 3m47s. The dispatch was backend=all, tests=prewarm,search, so the leg searched a near-empty database: 2.2M requests in 136 s, which made a 9.17 GB (8.5 GiB), 26.3M-line search-points.json. The runner itself survived. The || fallback printed the warning, and the job then ended. ci(bench): harden shared-host cleanup in fhir-benchmark.yml #1543's teardown ordering had already stopped the containers.

In the two earlier runs, the job was cancelled in this step and every later step was skipped.

Every new leg added by #1475 runs this step, which makes it more likely to hit.

Fix

  • New .github/scripts/fhir-bench/search_by_shape.py reads the file line by line and keeps only per-shape duration arrays (array('d')). It produces the same table as the jq program:
    • jq's //: null and false count as missing, while an empty string does not;
    • jq's round, half away from zero;
    • the same median, p95 and max indices;
    • sort_by(-.p95) with ties in ascending shape order, as group_by gives.
  • The step calls it with the same fallback warning. It gets timeout-minutes: 15, and matrix.backend moves to env: so the run block holds no expressions.

Verification

  • Equivalence with the old jq program, extracted verbatim from the workflow and run with jq 1.8.2:

    • Mixed fixture (ties, missing/null/false/empty tags, non-Point lines, other metrics, n=1/n=2, .5 values): byte-identical.
    • Empty file: identical (the header lines only).
    • With one malformed line: jq exits 5 and prints nothing. The script skips the line and prints the same table.
    • 3,000,000-line fixture (365 MB, 300k matching points across 20 shapes): byte-identical output.
    Wall time Peak memory
    search_by_shape.py ~5 s ~16 MB
    jq -rs ~68 s ~4.7 GB
  • actionlint 1.7.12: clean.

  • Live run: the same dispatch that lost the runner (-f backend=all -f tests=prewarm,search) is re-running from this branch as 36427203450. Results will be added here.

Refs #1475

🤖 Generated with Claude Code

https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ

…g it

"Attribute search latency per query shape" ran `jq -rs` over the k6
points file, which holds every metric of every search request. `-s`
loads the whole stream into one in-memory array, so memory grows with
the file, not with the handful of query shapes being summarised. The
self-hosted runner has been lost in this step three times: runs
30550776427, 33515369645 and 36423973848. In the last one, sqlite with
tests=prewarm,search searched a near-empty database very fast, and the
points artifact alone was 317 MB compressed.

.github/scripts/fhir-bench/search_by_shape.py reads the file line by
line and keeps only per-shape duration arrays. It produces the same
table: jq's `//` (null and false count as missing), jq's round-half-
away-from-zero, the same median/p95/max indices, and ties in ascending
shape order. On a 3,000,000-line fixture its output was byte-identical
to jq's: ~5 s and ~16 MB peak, against ~68 s and ~4.7 GB for jq -rs.
Unlike jq it also skips a malformed line instead of producing nothing.

The step also gets timeout-minutes: 15, and matrix.backend moves to
its env: so the run block holds no expressions.

Refs #1475

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013NudzWDu2yTGExYaxdTQYJ

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant