Conversation
Parse JSON documents into a tree borrowing strings from the input buffer and write them straight into the tantivy document, instead of building a serde_json::Map and tantivy OwnedValues first. The new path is used for JSON input without VRL transform and fingerprinter. It produces exactly the same partitions, documents and errors as doc_from_json_obj, which is checked by differential tests.
Make the fingerprinter generic over a small read-only JSON view implemented for serde_json values and BorrowedJsonDoc, so the DocProcessor can use the borrowed fast path when document clustering is enabled. Hashing is unchanged. Differential tests check that both views produce the same fingerprints.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
The DocProcessor is the bottleneck of a single indexing pipeline: one thread at 100%. About 45% of its time goes to parsing JSON into a
serde_json::Map(BTreeMap inserts, string allocations), and about 50% goes to converting that map intoOwnedValues and then into the tantivy document (copies, drops).This PR adds a second conversion path that avoids those intermediate owned values:
BorrowedJsonDoc::parsebuilds a small JSON tree. Strings are borrowed from the input buffer unless they contain escapes, and UTF-8 is validated once per document.DocMapper::doc_from_borrowed_jsonwalks the mapping tree over that tree and writes values directly into the tantivy document. Objects (_source, json fields, the dynamic field) are exposed to tantivy through itsValuetrait instead of being converted toOwnedValue.The DocProcessor uses this path for JSON input when no VRL transform is configured. Every other case (VRL, OTLP, plain text) still uses the existing
doc_from_json_objpath, which is unchanged apart from moving its shared tail into a helper.Document clustering. The
Fingerprinteris now generic over a small read-only JSON view, implemented for&serde_json::ValueandBorrowedJsonDoc. The view needs only four operations: value kind, array and object iteration through callbacks (no allocation), and key lookup. The hashing code is unchanged, sofingerprint()returns the same values as before. The newfingerprint_borrowed()returns the same fingerprint for the same JSON input.Equivalence with the existing path
The new path must produce exactly the same partition, the same field values in the same order, and the same errors as
doc_from_json_obj. To do so:serde_json::Map. By default that means sorted keys, with the last duplicate winning.preserve_orderis detected at runtime. That serde_json feature is enabled by a test dependency under--all-features(throughaws-smithy-http-client/test-util), so the parser checks the actualMapbehavior and mirrors it. It also mirrors the top-level key sort thatadd_objectapplies through itsBTreeMap.serde_json::Number. Coercion and partition hashing therefore stay identical.serde_json::ValuetoOwnedValue, and the new path does the same.There is one documented difference. A nested object whose first key is serde_json's private
$serde_json::private::RawValuetoken is rejected. The old path parsed its value as embedded JSON.Tests
serde_jsonon valid and invalid inputs, including duplicate keys, escapes, big and negative numbers, non-object roots, and invalid UTF-8.serde_json::Mapimplementation for several routing expressions._source, the default mapping, and an OTel-like mapping). The comparison checks ordered field values, partitions and errors.serde_json/preserve_order.max_tokens) policies. The DocProcessor test now checks that the fingerprint is equal to the owned one on a document with a duplicate key.quickwit-doc-mapperbehind thetestsuitefeature, so both crates use it.Note:
test_concatenate_multiple_fieldalready fails onmainwhen run without--all-features, because its expected order assumespreserve_order. This PR does not change that test.Performance
Micro-benchmark: 300 real OTel log rows, ~1.2 KB/doc, conversion only.
doc_from_json_bytes(before)doc_from_borrowed_json(after)End-to-end: Parquet local ingest of OTel logs, measured with this PR applied on top of #6843 (where Parquet rows also go through this path). Validation disabled, steady state over 40 s, two alternating runs per configuration on an M3 Max.
With document clustering enabled (policies:
structure, rawServiceName, rawSeverityText, tokenizedBody):With clustering, fingerprinting is about 40% of DocProcessor time, mostly in the
structurepolicy. That will be addressed in a separate PR.Parquet loads (#6843) use this path with no extra change: the pipeline creates the doc processor with the
jsoninput format.How was this PR tested?
cargo nextest run -p quickwit-doc-mapper, with and without--features serde_json/preserve_ordercargo nextest run -p quickwit-indexing doc_processorcargo clippy,make fmt