Skip to content

perf(lmdb): batch radix DFS reads + cache probe DBI handle - #11

Closed
hungpham10 wants to merge 61 commits into
Cleboost:mainfrom
hungpham10:perf/lmdb-batch-read
Closed

hungpham10 wants to merge 61 commits into
Cleboost:mainfrom
hungpham10:perf/lmdb-batch-read

Conversation

@hungpham10

Copy link
Copy Markdown

Vấn đề

Search song song trên LMDB chạm giới hạn reader-slot → MDB_BAD_RSLOT.

Nguyên nhân: mỗi get_node/get_children là một begin_ro_txn riêng. Một lần
DFS duyệt ~N node thì tạo ~2N read-txn. Nhiều request chạy song song (MCP/GraphQL
dùng chung Arc<GraphIndex>) thì số read-txn đồng thời nhân lên.

Thêm nữa, SharedGraphIndex::ensure_fresh → current_version → probe_version
chạy trước mọi request, mỗi lần lại open_db + begin_ro_txn.

Thay đổi

Tầng 1 — cache DBI handle trong probe_version

probe_env giữ thêm Database handle (Copy) cho sg_meta, thay vì open_db
lại mỗi request. Mỗi request còn đúng 1 read-txn + 1 get.

Tầng 2 — batch read trên CategoryStorage

Thêm get_nodes(&[usize]) / get_childrens(&[usize]) đọc nhiều id trong một
read-txn. Default impl gọi lại hàm đơn lẻ nên 6 backend nền (in-memory, sqlite,
redis, postgres, mysql, cached) không đổi. LMDB override.

Radix::search_dfs gom các lần đọc: đọc node của frame + toàn bộ child chưa duyệt
trong một lệnh gọi, thay vì một lệnh gọi cho từng node. Số read-txn mỗi vòng DFS
giảm từ ~2×children xuống 1. Biến base giữ nguyên thứ tự duyệt nên hành vi DFS
không đổi.

Test

  • batch khớp get_node/get_children từng cái — thứ tự và sort giữ nguyên
  • id không tồn tại vẫn trả BranchOutOfRange như bản đơn lẻ
  • input rỗng trả rỗng, không mở txn
  • regression: 8 task đọc song song (~800k lượt đọc) không lỗi
  • probe_version qua cache handle 100 lần vẫn đúng version

Benchmark

lmdb_batch_read (feature lmdb) — đọc 1/16/64 node đơn lẻ vs batch, cả với
get_children:

cargo bench -p codegraph-graph --bench lmdb_batch_read --features lmdb

Ghi chú

  • Không đụng write path.
  • Không thêm API _bulk song song với bản cũ — thay đổi input trực tiếp trên
    trait, backend nền dùng default impl.

hungpham7-tiki and others added 30 commits August 2, 2026 20:49
…-index-algorithm

Replace bfs with search index algorithm
* Remove old data structure

* Benchmark

* Remove temporary codegraph-viz and move benches into .github

* Fix issue codspeed

* Fix benchmark

* Implement sandbox and diff

* Add sandbox execution for diff and origin

* Remove unused unittest

* Fix lint

* Remove unused unittest

* Temporal remove another environments
* Implement new storage to improve performance

* style: apply rustfmt

* Fix lint

* Fix issue multiple reading in multiple streams
* Finetune to reduce LLM token

* style: apply rustfmt

* Fix lint
…es (#6)

* Implement MCP server for GA

* Setup unit-tests to new storages

* style: apply rustfmt

* Potential fix for pull request finding 'CodeQL / Workflow does not contain permissions'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* Add tests

* Fix tests

* Fix tests

* Fix tests

* Fix lint

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* Implement new embed vector and similar search

* style: apply rustfmt

* Fix unit test

* Add badges

* Fix coverage
Add a security policy document outlining supported versions and vulnerability reporting.
* Truncate APIs and support graphql with mermaid

* style: apply rustfmt
* Truncate APIs and support graphql with mermaid

* Implement to support windows

* style: apply rustfmt

* Potential fix for pull request finding 'CodeQL / Workflow does not contain permissions'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* Fix lint

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* Implement flow to install codegraph to MacOS and linux

* Add missing pipeline to support releasing

* Fix issue with doctor

* style: apply rustfmt
…rce code as possible (#13)

* Bump version to v2.0.3

* Update ci

* Update ci again

* Remove coverage

* Setup codecov.yml
hungpham10 and others added 26 commits August 31, 2026 07:56
* Improve code to be more maintainable

* style: apply rustfmt

* Fix lint
- Add support to build graph from binary using radare2 for reverse binary
code
- Fix issue when working with lambda function which is a core feature to
analyze obfuscated code
* Fix issue missing function name and crash when picking to wrong function

* Bump version to v2.1.1

* style: apply rustfmt
* Improve by showing more meaningful object naming from r2

* Increase code-coverage

* style: apply rustfmt
* Implement tool to convert structure documents into graph

* Bump version to v2.1.3
* Integrate codegraph-docs into main flow

* style: apply rustfmt

* Fix lint

* style: apply rustfmt
* Persistent document graph into disk

* Inplement nginx parser

* style: apply rustfmt

* Fix lint

* style: apply rustfmt

* Bump version to v2.1.5
* Split storage of binary graph into different database

* Fix lint
* Support YAML multiple files

* style: apply rustfmt

* Bump version to v2.1.6
* Show progress bar while indexing documents

* Improve performance when starting

* Bump version to v2.1.7

* style: apply rustfmt

* Fix lint
* Fix document graph pipeline and wire real pattern/value search

The document tools were unusable end-to-end due to four latent bugs:

- parsers: parent->children links were never persisted (nodes cloned into
  Document.nodes before children wiring), so hydrate could never descend.
- graph: path token chains were built in reverse order (root ended up last),
  so full-path queries never matched; also node cache was materialized after
  trie insertion, so the first ingest interned no keys.
- graph: the four trie projections shared one storage without namespace —
  radix root pointers collided per shard, leaving only the last-written trie
  reachable. Add shard_bias to Radix/Search (default 0) and give each docs
  trie a distinct shard range.
- graph: the string interner was RAM-only while tries persisted, so interned
  token payloads dangled after restart. Persist the interner blob alongside
  docs and restore it in open(); doc ids move to their own id range (>=6e11)
  so they no longer collide with node ids.

MCP/CLI wiring:
- doc_search now parses dotted patterns into token chains via the interner,
  with a case-insensitive key-scan fallback for non full-path queries.
- new tools: doc_search_value (scalar substring scan), doc_ingest_dir
  (recursive bulk ingest with limit), doc_remove.
- doc_hydrate takes max_depth to keep LLM payloads small; doc_list returns
  per-doc metadata (doc_id, path, format, root_node_id, nodes) instead of
  bare counts; search results include key/index.
- CLI: fix `doc ingest` clap panic (positional `path` clashed with the
  global --path arg, renamed to `file`); doc search uses the same real
  pattern resolution.

* Bump version to v2.1.8

* style: apply rustfmt
#30)

* Add fuzzy key match, pattern mining (P#) and IDF ranking to document graph

Phase 3 of the structured document graph: fuzzy retrieval, cross-document
pattern mining and rarity-based ranking (rare structure = high information,
IDF-style scoring).

- Fuzzy key match: search_key_fuzzy scores distinct keys by exact/prefix/
  contains/Levenshtein similarity with an IDF bonus (rarer keys score
  higher). `doc_search` accepts `~segment` to force fuzzy and falls back
  exact-substring -> fuzzy when the full-path trie misses.
- Pattern mining: mine_patterns() counts kind chains (wildcard FIELD/IDX
  payloads) ending at scalar leaves over a max_depth window, assigns stable
  pattern ids (P#, registry persisted in storage and restored on open) and
  indexes chains into the previously dead pattern_trie. Results carry
  node_count / doc_count / doc_freq, sorted by document frequency ascending
  so characteristic patterns surface first and background (~1.0) sinks.
- Structural search: search_kind_chain() matches nodes whose ancestor kind
  window ends with the query chain (e.g. "MAP, FIELD, NUMBER"); results are
  ranked by pattern uniqueness IDF. search_path_scan() complements the radix
  trie whose leaf holds a single record per chain, returning all nodes with
  the same key path across documents (spec.replicas: 97 hits on the infra
  repo instead of 1).
- Depth semantics fixed in doc_search: depth counts extra levels BELOW the
  pattern, so the radix key-length filter gets pattern_len + depth.
- New MCP tools: doc_mine_patterns, doc_list_patterns, doc_search_struct.
  New CLI commands: `codegraph doc patterns`, `codegraph doc struct`.

* style: apply rustfmt

* Fix lint

* Bump to version v2.1.9
#31)

* Fix binary callees/callers/flow always empty: route id >= bin_base to BinaryGraph (#26 regression)

After the storage split, binary symbols live only in binary.sqlite (id range
bin_base = 2e9) while codegraph_callees/callers/flow still queried the main
GraphIndex, so every binary returned empty results. Now those tools route to
BinaryGraph when the node id falls in the binary range, with a fallback to the
old path if the binary DB is unavailable.

- BinaryGraph: add callees (call records resolved by name within the same
  binary), callers (BFS over a new caller_of:{name} reverse index built at
  ingest) and flow (chain render with CFG markers + call sites, same shape as
  GraphIndex::flow).
- Ingest: stop corrupting chain entries — CFG markers (id < SYMBOL_BASE) and
  unresolved-call placeholders (0) are kept as-is; only real symbol ids are
  remapped into the bin_base range.
- MCP: dispatch_binary_graph routes callees/callers/impact/flow for binary ids.
- Config template: document the [bingraph] section (enabled/bin_base/storage)
  with a warning against overlapping id ranges.
- Docs: binary-analysis.md updated for the split-storage architecture.
- Tests: bingraph unit tests extended + end-to-end test compiling a real
  shared library and running it through r2 extraction and queries.

* style: apply rustfmt
r2 sometimes fails to recover call ops for a shared library (entrypoint
detection fails on Mach-O dylibs), producing chains without call sites. Retry
extraction up to 3 times with the extract cache disabled and only accept a run
whose chains contain at least one resolved call.
- codegraph_graphcode_*: class, list_types, function_scope,
  search_by_annotation, files, dependencies, sandbox, diff,
  diff_simulate, origin_simulate + new codegraph_graphcode_stats
- codegraph_doc_* -> codegraph_graphdoc_* (all 11 tools)
- codegraph_binary_list/addr/stats -> codegraph_graphbin_*;
  binary_search merged into codegraph_search_symbol via new
  `source` arg (all|code|binary)
- Shared/session tools keep their names (symbol, search_symbol,
  callers, callees, impact, flow, context, references, mermaid,
  search_flow, init/deinit/index, query_usage_report)
- codegraph_status now aggregates all three datasets (code + doc +
  binary, null when absent)
- Minimize everywhere: doc/binary outputs routed through
  emit_value (omit_defaults); graphbin list/addr support
  `format` with fixed-order row arrays
- Docs updated: server-instructions.md, README, architecture,
  binary-analysis
* Update graphql to support new models

* Bump version to v2.2.1
* Add subcommand `clean` and remove subcommand `doc`

* Bump version to v2.2.2
* Add resume-able functions to avoid hanging

* Optimize cost by compressing response from MCP

* Fix lint

* Change minimize to minimal
Search song song tren LMDB chạm gioi han reader-slot (`MDB_BAD_RSLOT`):
moi `get_node`/`get_children` la mot `begin_ro_txn` rieng, mot lan search
tao O(nodes) read-txn, nhieu request thi nhan lai gap gioi han.

Tang 1 - cache DBI handle trong probe_version:
`probe_env` giu them `Database` handle (Copy) cho `sg_meta` thay vi
`open_db` lai moi request. `SharedGraphIndex::ensure_fresh` chay
`current_version` truoc moi request nen moi lan deu ton mot `open_db`
(tha bang `Environment`) va mot `begin_ro_txn`; gio con 1 read-txn.

Tang 2 - batch read tren CategoryStorage:
them `get_nodes`/`get_childrens` doc nhieu id trong MOT read-txn. Default
impl goi lai ham don le nen 6 backend nen (in-memory/sqlite/redis/pg/mysql/
cached) khong doi. LMDB override de dung 1 txn cho ca danh sach.

`Radix::search_dfs` chuyen sang goy cac lan doc: doc node cua frame + tat ca
child chua duyet trong 1 lenh goi, thay vi mot lenh goi cho tung node. So
read-txn moi vong DFS giam tu ~2*children xuong 1. `base` giu nguyen thu tu
duyet nen hanh vi DFS khong doi.

Test:
- batch khop `get_node`/`get_children` tuang tung, thu tu va sort giu nguyen
- id khong ton tai van tra `BranchOutOfRange` nhu ban don le
- regression: 8 task doc song song khong loi
- `probe_version` qua cache handle 100 lan van dung version

Benchmark moi `lmdb_batch_read` (feature `lmdb`) do doc 1/16/64 node don
lue vs batch, ca voi `get_children`:
cargo bench -p codegraph-graph --bench lmdb_batch_read --features lmdb

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@hungpham10 hungpham10 closed this Oct 3, 2026
@hungpham10
hungpham10 deleted the perf/lmdb-batch-read branch October 3, 2026 03:40
@hungpham10
hungpham10 restored the perf/lmdb-batch-read branch October 3, 2026 03:45
@Cleboost

Cleboost commented Oct 4, 2026

Copy link
Copy Markdown
Owner

bro? wtf?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants