Skip to content

Faster JSON doc conversion in the DocProcessor - #6866

Draft
fmassot wants to merge 2 commits into
mainfrom
fmassot/borrowed-doc-parsing
Draft

fmassot wants to merge 2 commits into
mainfrom
fmassot/borrowed-doc-parsing

Conversation

@fmassot

@fmassot fmassot commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Description

The DocProcessor is the bottleneck of a single indexing pipeline: one thread at 100%. About 45% of its time goes to parsing JSON into a serde_json::Map (BTreeMap inserts, string allocations), and about 50% goes to converting that map into OwnedValues and then into the tantivy document (copies, drops).

This PR adds a second conversion path that avoids those intermediate owned values:

  1. BorrowedJsonDoc::parse builds a small JSON tree. Strings are borrowed from the input buffer unless they contain escapes, and UTF-8 is validated once per document.
  2. DocMapper::doc_from_borrowed_json walks the mapping tree over that tree and writes values directly into the tantivy document. Objects (_source, json fields, the dynamic field) are exposed to tantivy through its Value trait instead of being converted to OwnedValue.

The DocProcessor uses this path for JSON input when no VRL transform is configured. Every other case (VRL, OTLP, plain text) still uses the existing doc_from_json_obj path, which is unchanged apart from moving its shared tail into a helper.

Document clustering. The Fingerprinter is now generic over a small read-only JSON view, implemented for &serde_json::Value and BorrowedJsonDoc. The view needs only four operations: value kind, array and object iteration through callbacks (no allocation), and key lookup. The hashing code is unchanged, so fingerprint() returns the same values as before. The new fingerprint_borrowed() returns the same fingerprint for the same JSON input.

Equivalence with the existing path

The new path must produce exactly the same partition, the same field values in the same order, and the same errors as doc_from_json_obj. To do so:

  • Objects are ordered and deduplicated like serde_json::Map. By default that means sorted keys, with the last duplicate winning.
  • preserve_order is detected at runtime. That serde_json feature is enabled by a test dependency under --all-features (through aws-smithy-http-client/test-util), so the parser checks the actual Map behavior and mirrors it. It also mirrors the top-level key sort that add_object applies through its BTreeMap.
  • Numbers are stored as serde_json::Number. Coercion and partition hashing therefore stay identical.
  • RFC 3339 detection is mirrored. tantivy turns strings that look like RFC 3339 dates into dates when converting serde_json::Value to OwnedValue, and the new path does the same.
  • Error messages are reproduced. Each mirrored function documents which function it mirrors.

There is one documented difference. A nested object whose first key is serde_json's private $serde_json::private::RawValue token is rejected. The old path parsed its value as embedded JSON.

Tests

  • Parser: the tree is compared against serde_json on valid and invalid inputs, including duplicate keys, escapes, big and negative numbers, non-object roots, and invalid UTF-8.
  • Partition hashing: compared against the serde_json::Map implementation for several routing expressions.
  • Differential tests: old and new paths are compared on 8 doc mappings (dynamic, lenient, strict, with and without _source, the default mapping, and an OTel-like mapping). The comparison checks ordered field values, partitions and errors.
    • Inputs are ~45 hand-written documents, including many error cases, plus 20,000 seeded random documents.
    • The suite catches each of five mutations I tried, for example removing date detection, keeping the first duplicate, or reordering number classification.
    • The suite passes with and without serde_json/preserve_order.
  • Fingerprints: owned and borrowed fingerprints are compared on 13 hand-written and 20,000 random documents. The config uses structure (with and without exclusions), raw (nested paths, objects, arrays) and tokenized (with max_tokens) policies. The DocProcessor test now checks that the fingerprint is equal to the owned one on a document with a duplicate key.
  • The random document generator lives in quickwit-doc-mapper behind the testsuite feature, so both crates use it.

Note: test_concatenate_multiple_field already fails on main when run without --all-features, because its expected order assumes preserve_order. This PR does not change that test.

Performance

Micro-benchmark: 300 real OTel log rows, ~1.2 KB/doc, conversion only.

µs/doc
doc_from_json_bytes (before) 6.4
parse + doc_from_borrowed_json (after) 3.45

End-to-end: Parquet local ingest of OTel logs, measured with this PR applied on top of #6843 (where Parquet rows also go through this path). Validation disabled, steady state over 40 s, two alternating runs per configuration on an M3 Max.

Pipelines Before (docs/s) After (docs/s) Change
1 208k 322k +54%
8 623k–688k 863k–915k +36%

With document clustering enabled (policies: structure, raw ServiceName, raw SeverityText, tokenized Body):

Pipelines Before (docs/s) After (docs/s) Change
1 151k 205k–207k +37%
8 498k–515k 563k–594k +15%

With clustering, fingerprinting is about 40% of DocProcessor time, mostly in the structure policy. That will be addressed in a separate PR.

Parquet loads (#6843) use this path with no extra change: the pipeline creates the doc processor with the json input format.

How was this PR tested?

  • cargo nextest run -p quickwit-doc-mapper, with and without --features serde_json/preserve_order
  • cargo nextest run -p quickwit-indexing doc_processor
  • cargo clippy, make fmt
  • End-to-end Parquet benchmark as described above

Parse JSON documents into a tree borrowing strings from the input buffer and
write them straight into the tantivy document, instead of building a
serde_json::Map and tantivy OwnedValues first.

The new path is used for JSON input without VRL transform and fingerprinter.
It produces exactly the same partitions, documents and errors as
doc_from_json_obj, which is checked by differential tests.
Make the fingerprinter generic over a small read-only JSON view implemented
for serde_json values and BorrowedJsonDoc, so the DocProcessor can use the
borrowed fast path when document clustering is enabled.

Hashing is unchanged. Differential tests check that both views produce the
same fingerprints.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant