Conversation
7e71326 to
a0bcbe1
Compare
d5ec73f to
0277a97
Compare
Feature-gated, benchmark-only local Parquet ingestion with parallel indexing pipelines, bypassing the ingest API and WAL. CLI-injected sources share a row-group plan and emit JSON documents; servers keep the default source loader. No merge pipeline is spawned. Loads require an empty index, reject invalid documents, and attempt index cleanup on failure. Uploader lifecycle changes are kept in a separate independent PR.
0277a97 to
70b38e2
Compare
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 70b38e24a2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| itertools = { workspace = true } | ||
| libz-sys = { workspace = true, optional = true } | ||
| oneshot = { workspace = true } | ||
| parquet = { workspace = true, optional = true, features = ["arrow", "snap", "zstd", "lz4"] } |
There was a problem hiding this comment.
Enable standard GZIP and Brotli Parquet codecs
When the input file uses GZIP or Brotli compression—both standard Parquet codecs—the reader will reject the file because default-features = false disables those decoders and this feature list enables only Snappy, Zstd, and LZ4. Since the CLI and documentation advertise generic .parquet input without a codec restriction, enable the flate2 and brotli Parquet features (or explicitly reject/document those otherwise-valid files).
Useful? React with 👍 / 👎.
Motivation
There is no good way to index very fast data from parquet files (or equivalent well compressed file format). This can really help some use cases like large migrations or benchmark.
Solution
Add a CLI command to bulk load of a local Parquet file with N indexing pipelines, bypassing the ingest API and WAL:
Behind the
parquetCargo feature, off by default. No config or proto change.Details:
docs/internals/parquet-bulk-load.md.Tests
309 CLI/indexing tests passed; workspace clippy with and without Parquet, formatting and license checks passed.
Used for the https://github.com/serenedb/searchbench and managed to divide by >2 the time to index.