Skip to content

search: chained and _has searches whose resolved id set exceeds 10,000 answer 500 on the Elasticsearch composites (max_result_window) #1548

Description

@mauripunzueta

Summary

On the Elasticsearch composites, a chained or _has search whose resolved id set exceeds Elasticsearch's index.max_result_window (10,000) answers 500 with an opaque OperationOutcome (code: exception, "An internal error occurred"). Two rows of the manual matrix (MANUAL_TESTING_MATRIX.md §8.2, rows 4.10 and 4.11) fail this way on mongo-es with the full Synthea corpus.

Current behavior

main at 722a0927b, HFS_STORAGE_BACKEND=mongo-es, 18.96 M resources (7.7 M Observations, 11,705 Patients), Elasticsearch 8.15.0 with the default max_result_window:

Matrix row Request Result
4.10 chained GET /Observation?subject:Patient.family=Parker433&_count=5 500
4.10 chained (untyped) GET /Observation?subject.family=Parker433&_summary=count 500
4.11 reverse chained GET /Patient?_has:Observation:patient:code=http://loinc.org|8302-2&_count=5 500
control GET /Observation?subject.identifier=http://hl7.org/fhir/sid/us-ssn|999-33-3920 200, 165
control GET /Patient?_has:Condition:patient:code=http://snomed.info/sct|706893006 200, 5,294

Server log for each failure:

WARN helios_rest::handlers::search: Chained search resolution failed error=internal error in elasticsearch:
Search failed after 3 attempts (status 400): {"error":{"root_cause":[{"type":"illegal_argument_exception",
"reason":"Result window is too large, from + size must be less than or equal to: [10000] but was [10001].
See the scroll api for a more efficient way to request large data sets. This limit can be set by changing
the [index.max_result_window] index level setting."}], ... "phase":"query"
ERROR helios_rest::error: internal error while processing request error.detail=internal error in elasticsearch: ...
ERROR tower_http::trace::on_failure: response failed classification=Status code: 500 Internal Server Error latency=2198 ms

The controls pass because their resolved sets are small (165 Observations; 5,294 Patients). The failures resolve more than 10,000 ids: the 30 Parker433 patients own roughly 20,000 Observations, and 8302-2 matches 175,355 Observations.

Root cause

crates/persistence/src/search/chain_resolver.rs materialises every hop of a chain as a full id list with offset paging (search_all_pages, RESOLVER_PAGE = 1000, page.offset = Some(offset) until a short page). The Elasticsearch backend clamps size so from + size never exceeds max_result_window (search/query_builder.rs, #1079), so the resolver's eleventh page asks for from=10000, size=1 and Elasticsearch rejects it. The error is a plain BackendError, which the REST layer maps to 500.

The SQL and MongoDB backends have no such window, so the same searches succeed (slowly) there; the composites sqlite-es, pg-es, mongo-es and s3-es all route search to Elasticsearch and are all affected. ElasticsearchConfig::max_result_window is not exposed as an environment variable, so there is no deployment-side workaround.

Expected behavior

Either the chain resolves regardless of size (page the intermediate hops with search_after/PIT, or resolve the terminal hop inside Elasticsearch with a terms lookup instead of materialising ids), or the server answers a 4xx that names the limit (too-costly) instead of a 500. A chained search that succeeds on sqlite should not be an internal error on sqlite-es.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions