Hybrid semantic search for a local document corpus. Combines BM25 (Recoll/Xapian), vector search (BGE-M3 + Cohere + ChromaDB HNSW), 5-arm Reciprocal Rank Fusion, HyDE query expansion (optional), and Claude agents for ranked, cited answers.
Tested on a 144 GB multilingual corpus — 128K files, 950K chunks — PDFs, Word files, spreadsheets, scanned images (with OCR), emails, and plain text in Greek + English.
Each document is assigned an authority tier (A = official/institutional → D = noise) at index time via a regex classifier with optional Haiku LLM fallback. Tier acts as a tiebreaker in RRF ranking and surfaces as citation labels in answers (⚠ unconfirmed for C-tier personal notes).
Current performance (v7.2, production): 54.3% Recall@20 on a 70-query stress-test golden set (38 easy + 32 hard queries). Easy subset: 78.9% (30/38). Note: the older 70% figure was on a different, easier 105-query eval set.
dsearch "contract deadline extension" # fast terminal results
dsearch -w "invoice Q3" # sortable HTML table in browser
dsearch -CLexpand-CLrank-CLanswer "my query" # full AI answer with citations (~4 min)
dsearch -CLexpand-CLrank-CLanswer -deep "query" # deep multi-hop retrieval (~7 min)
dsearch-multimodel --model gte "my query" # compare across 5 embedding models
dsearch2 "pension fund certificate" # v2: hybrid vector+TF with RRF merge
dsearch2 --json "insurance claim" | jq . # structured JSON output
dsearch2 --debug "ασφαλιστικό" # show alias expansion + fetch stats
Typical cost per answer run: ~$0.01–0.02 (Cohere API only — Claude runs via your existing subscription, no per-token billing).
Deep mode: ~$0.02–0.04 (a few extra Cohere queries for the extra retrieval rounds).
Interactive diagrams: Architecture — runtime pipeline (drag nodes, click for detail) · Setup guide — step-by-step with full shell commands · Full docs
Five models are indexed. dsearch-multimodel can search any of them independently for comparison.
| Key | Model | Dim | Notes |
|---|---|---|---|
bge-m3 ★ |
BAAI/bge-m3 |
1024 | Primary vector arm in dsearch2; strong MTEB benchmarks |
cohere-v3 |
cohere/embed-multilingual-v3.0 |
1024 | Optional (--cohere flag); best multilingual precision |
e5-base |
intfloat/multilingual-e5-base |
768 | Good retrieval asymmetry for long docs |
e5-large |
intfloat/multilingual-e5-large |
1024 | Broadest coverage for legal/procedural text |
gte |
Alibaba-NLP/gte-multilingual-base |
768 | Sharpest score distribution; 8 K context |
★ = primary (dsearch2 default)
System packages
# Debian/Ubuntu
sudo apt install python3 python3-venv python3-pip \
recoll recollindex \
poppler-utils antiword libreoffice-common- Recoll — BM25 full-text indexer (provides
recollindex+recollq) - poppler-utils —
pdftotextfor PDF extraction - antiword —
.doc(legacy Word) extraction - LibreOffice —
.docx/.odtfallback extraction
Claude CLI — the answer pipeline spawns Claude via subprocess:
npm install -g @anthropic-ai/claude-code # or follow https://claude.ai/code
claude --version # verify it worksAPI keys
- Cohere API key — free tier works for queries; paid needed for large index builds
- Anthropic subscription for Claude CLI (used by
dsearch-answer)
git clone https://github.com/gruncode/semantic-disk-search.git
cd semantic-disk-searchTwo venvs are required because GTE needs transformers<=4.49 while other models need newer versions.
# Venv 1 — used by dsearch CLI, Cohere, ChromaDB, GTE model
python3 -m venv ~/venvs/gte-embed
~/venvs/gte-embed/bin/pip install \
cohere chromadb faiss-cpu numpy \
"transformers==4.49" torch \
sentence-transformers \
json-repair regex \
openpyxl xlrd python-pptx
# Venv 2 — used by e5, bge-m3, Jupyter notebooks
python3 -m venv ~/venvs/transformers
~/venvs/transformers/bin/pip install \
faiss-cpu numpy sentence-transformers \
"transformers>=5.0" torch jupytermkdir -p ~/.config/dsearch
cp .env.example ~/.config/dsearch/.env
$EDITOR ~/.config/dsearch/.envFill in at minimum:
COHERE_API_KEY=your-key-here
CORPUS_DIR=~/Documents # root of what you want to search
FAISS_BASE=~/.local/share/dsearch/faiss
CHROMADB_DIR=~/.local/share/dsearch/chromadb
XAPIAN_DB=~/.local/share/dsearch/xapiandb
VENV_GTE=~/venvs/gte-embed/bin/python3
VENV_TF=~/venvs/transformers/bin/python3
HF_HOME=~/.cache/huggingfacesudo cp scripts/dsearch scripts/dsearch-answer scripts/dsearch-multimodel /usr/local/bin/
sudo chmod +x /usr/local/bin/dsearch /usr/local/bin/dsearch-answer /usr/local/bin/dsearch-multimodelYou need to build two indexes: a BM25 index (Recoll) and a vector index (ChromaDB or FAISS). Both read from $CORPUS_DIR.
# Configure Recoll to index your corpus
mkdir -p ~/.recoll
echo "topdirs = $HOME/Documents" >> ~/.recoll/recoll.conf
# Run the indexer (takes 30 min – several hours depending on corpus size)
make recoll-reindex
# Check progress
tail -f /tmp/recoll-reindex.logThis is what dsearch uses by default. Costs ~$1.60 per million chunks via the Cohere API.
# Step 1: extract text chunks from all files → chunks.jsonl
$VENV_GTE src/extract_chunks.py $CORPUS_DIR \
--output /tmp/chunks.jsonl \
--chunk-size 500 --overlap 100
# Step 2: embed with Cohere → embeddings.npy
set -a && source ~/.config/dsearch/.env && set +a
$VENV_GTE scripts/cohere_embed.py \
--chunks /tmp/chunks.jsonl \
--output /tmp/embeddings.npy
# Step 3: load into ChromaDB
$VENV_GTE scripts/build_chroma_from_embeddings.py \
--chunks /tmp/chunks.jsonl \
--embeddings /tmp/embeddings.npy \
--db $CHROMADB_DIR --collection fulldisk# GTE model (~6.5 h on CPU for a large corpus)
make rebuild-gte
# E5-base
make rebuild-e5dsearch "pension fund certificate"
dsearch -n 20 "water leak insurance claim" # top 20 resultsOutput: colour-coded ranked list (green = high score, yellow = medium, red = low) with file path, type, date, and a snippet with query words highlighted.
dsearch -w "building permit extension"Opens a sortable, filterable table in your browser. Click any column header to sort. Filter box matches across all columns. "Group by folder" clusters results from the same directory.
dsearch -CLexpand-CLrank-CLanswer "why was the project delivery delayed"Runs the full 6-step pipeline. Prints timestamped progress for each step, then a cited answer like:
ANSWER:
The project was delayed for three reasons:
1. Supplier lead time exceeded the contract deadline by 6 weeks [delivery-timeline.pdf]
2. Site access was restricted due to permit issues [permit-correspondence.docx]
3. ...
SOURCES:
[1] ~/Documents/Projects/delivery-timeline.pdf (score 9.2)
[2] ~/Documents/Legal/permit-correspondence.docx (score 8.7)
...
dsearch -CLexpand-CLrank-CLanswer -deep "comprehensive review of Siemens proposal"After the initial retrieval, Claude reads the top results, extracts new domain-specific terms, and runs up to 3 additional search rounds. Adds ~1–2 minutes but substantially improves recall on vocabulary-mismatch queries.
dsearch-multimodel "pump maintenance schedule" # default: cohere-v3
dsearch-multimodel --model gte "pump maintenance schedule"
dsearch-multimodel --model e5 "pump maintenance schedule"dsearch2 is a standalone hybrid retriever with a 5-arm architecture fused through weighted RRF (Reciprocal Rank Fusion): BGE-M3 vector (W=1.0), TF-IDF lexical (W=1.0), Recoll BM25 path (W=1.5), Recoll BM25 content (W=0.0, disabled), and FTS5 (W=0.0, disabled) — plus alias expansion for multilingual synonym matching and document tier classification (A/B/C/D reliability tiers via regex + optional LLM fallback). BGE-M3 is the primary vector arm; Cohere is retained as optional (--cohere to enable). Three-store architecture: Cohere ChromaDB + BGE ChromaDB + SQLite.
| Feature | dsearch | dsearch2 |
|---|---|---|
| Vector retrieval | Cohere + ChromaDB | BGE-M3 + ChromaDB (primary); Cohere optional |
| Lexical scoring | None | TF-IDF over returned chunks |
| BM25 integration | None | Recoll path + content (split channels) |
| Rank fusion | Vector rank only | 5-arm weighted RRF (BGE vector + TF-IDF + Recoll path + Recoll content + FTS5) |
| Alias expansion | Via Claude Step 1 | Built-in, from alias_map.json |
| Source tier | None | A/B/C/D classifier (regex + optional Haiku LLM) |
| Claude calls | Yes (Steps 1, 4, 6) | None — retrieval only |
| Output modes | terminal, HTML, answer | terminal, JSON |
| HNSW candidate pool | n_results × 3 | max(n_results × 3, 1000) |
| Recall@20 (70 stress-test queries) | — | 54.3% (v7.2) |
The larger candidate pool (max(..., 1000)) is critical for HNSW recall on large collections (950K+ chunks) where a small probe set degrades accuracy.
dsearch2 "pension fund certificate"
dsearch2 -n 30 "ασφαλιστικό ταμείο"
dsearch2 --json "invoice delay" | jq '.results[].path'
dsearch2 --show-aliases # list all alias groups
dsearch2 --debug "query" # print fetch count, alias expansion, top ranks
dsearch2 --no-alias "query" # disable alias expansionAll paths are read from ~/.config/dsearch/.env:
CHROMADB_DIR=~/.local/share/dsearch/chromadb
DSEARCH_COLLECTION=fulldisk-1k
DSEARCH_ALIAS_MAP=configs/alias_map.json
COHERE_API_KEY=your-key
# Optional: colon-separated paths to suppress from non-tool queries
DSEARCH_META_PREFIXES=/path/to/project-docsmake search Q="query" quick Cohere search
make search-gte Q="query" search with GTE model
make rebuild-gte rebuild GTE FAISS index (~6.5h CPU)
make rebuild-e5 rebuild e5-base FAISS index
make rebuild-cohere rebuild Cohere ChromaDB index (API cost)
make benchmark run evaluation against golden query set
make recoll-reindex full BM25 reindex (background)
make recoll-status show index sizes
make jupyter launch Jupyter notebook server
make status show all index sizes and doc counts
make help list all targets
scripts/
dsearch main CLI — terminal, HTML, and answer modes
dsearch-answer 6-step AI answer pipeline (called by dsearch)
dsearch-multimodel FAISS search across any of the 5 models
dsearch2 v2 hybrid retriever — vector + TF + RRF, alias expansion, tier boosts
cohere_embed.py batch embed chunks via Cohere API
build_chroma_from_embeddings.py load embeddings into ChromaDB
src/
extract_chunks.py parse corpus files → text chunks (one-time)
multi_search.py merge + dedup results from multiple models
recoll_vector_index.py build FAISS from Recoll Xapian text
chroma_vector_index.py ChromaDB index builder
recoll_query_assist.py BM25-only path (legacy, no vector)
recoll_benchmark.py evaluation harness (MRR, hit@k)
configs/
models.yaml model names, dimensions, index paths
alias_map.json bilingual query alias expansion table
docs/
index.html interactive architecture diagrams
flow/dsearch.html dsearch script flow diagram
flow/dsearch-answer.html answer pipeline flow diagram
recoll_experiments.md BM25 tuning experiment log
experiments-report.html full experiments report (HyDE, FTS5, weight sweeps)
PROJECT-REPORT.md full project report with benchmark results
notebooks/
01_model_comparison.ipynb 5-model benchmark comparison + charts
02_experiment_log.ipynb retrieval experiment history
03_pipeline_overview.ipynb pipeline architecture walkthrough
tests/
test_dsearch2_smoke.py basic CLI smoke tests
test_dsearch2_retrieval_eval.py retrieval quality tests
test_dsearch2_eval_tiers.py document tier classification tests
test_dsearch2_answer_diversity.py answer diversity tests
test_dsearch2_extraction_audit.py text extraction tests
test_source_tiers.py source reliability tests
Evaluated on a golden set of real queries across document types (contracts, invoices, technical specs, emails, scanned letters).
| Model | MRR | Hit@1 | Hit@5 |
|---|---|---|---|
| Cohere-v3 | 0.91 | 0.82 | 0.97 |
| E5-large | 0.87 | 0.76 | 0.95 |
| E5-base | 0.84 | 0.73 | 0.93 |
| GTE | 0.81 | 0.70 | 0.91 |
| BGE-M3 | 0.78 | 0.66 | 0.89 |
| BM25 only | 0.71 | 0.58 | 0.85 |
| BM25 + Cohere (hybrid) | 0.94 | 0.85 | 0.99 |
See notebooks/01_model_comparison.ipynb for the full comparison with charts.
| Version | N | Method | Recall@1 | Recall@20 |
|---|---|---|---|---|
| v5 | 105 | Vector only (Cohere) | 19% | 56% |
| v6 | 105 | Vector + Cohere Rerank | 29% | 58% |
| v7 | 105 | Vector + Recoll BM25 (equal weights) | 25% | 64% |
| v7.1 | 105 | Weighted RRF + split Recoll (path/content) | 29% | 69% |
| v7.1c | 105 | Grid-optimal weights + classifier + fixes | 29% | 70% |
| v7.2 | 70 | BGE-M3 + 5-arm RRF (stress-test set) | — | 54.3% |
v7.2 production weights: W_BGE_VEC=1.0, W_TF=1.0, W_RECOLL_PATH=1.5, W_RECOLL_CONTENT=0.0 (disabled), W_FTS5=0.0 (disabled)
Eval set change: v7.2 uses a harder 70-query golden set (38 easy + 32 hard stress-test queries). The 105-query set used through v7.1c was retired as too easy. Easy subset Recall@20: 78.9% (30/38).
Key insight: multi-signal RRF fusion outperforms a single reranker by a wide margin, demonstrating that complementary signals (vector + lexical + filesystem path BM25) provide more information than any single model.
BGE-M3 (1024-dim dense embeddings) is now fully integrated as the primary vector arm. During the pilot phase, A/B testing on 32 residual misses showed:
| Verdict | Count | Meaning |
|---|---|---|
| BGE-M3 wins | 10/32 | Target enters top-20 with BGE-M3 where Cohere failed |
| Cohere wins | 0/32 | BGE-M3 never loses to Cohere |
| Both hit | 2/32 | Both models find the target |
| Both miss | 20/32 | Neither finds it (OCR garbage, meaning-in-path only) |
Regression test on 73 queries showed 5/73 (6.8%) would regress on a full Cohere swap. This led to the dual-embed architecture: BGE-M3 as primary, Cohere retained as optional (--cohere). Three-store design: Cohere ChromaDB + BGE ChromaDB + SQLite.
SQLite FTS5 was tested as a 5th retrieval arm across 4 weight configurations. Result: the lexical axis is already saturated by TF-IDF + Recoll — FTS5 produced 0 gains and 1–2 regressions. Disabled (W=0.0).
No other personal disk search tool publishes Recall@K benchmarks. The closest comparisons are academic retrieval benchmarks:
| System / Benchmark | Recall metric | Score | Corpus |
|---|---|---|---|
| BM25 alone (BEIR avg) | nDCG@10 | ~40–45% | Clean English |
| Dense retrieval (BEIR SOTA) | nDCG@10 | 55–65% | Clean English |
| Hybrid BM25+dense (MS MARCO) | Recall@10 | 80.8% | Clean English |
| Cross-lingual hybrid (MKQA) | Recall@20 | 67.8% | Multilingual |
| Local hybrid RAG (tuned 30/70) | Recall@5 | 81.6% | Clean local docs |
| dsearch2 v7.1c (retired eval) | Recall@20 | 70% | Greek+English, OCR, mixed formats |
| dsearch2 v7.2 (stress-test) | Recall@20 | 54.3% | Greek+English, OCR, mixed formats (harder eval set) |
dsearch2 operates on a significantly harder corpus than standard benchmarks: bilingual (Greek + English), noisy OCR, meaning encoded in file paths/names, and extreme document diversity (payslips, legal docs, SCADA reports, teaching materials, dental X-rays).
Building FAISS indexes for large corpora on CPU takes many hours. You can use a cloud GPU instead:
# Start a GCP T4 spot instance
ssh <gcp-gateway> 'gcloud compute instances start $GCP_INSTANCE --zone=$GCP_ZONE'
# Transfer chunks, run embedding remotely
scp /tmp/chunks.jsonl gpu-instance:/tmp/
ssh gpu-instance 'python3 gpu_embed_multi.py --model gte'
# Copy indexes back
scp -r gpu-instance:/tmp/faiss-index/ $FAISS_BASE/
# Stop instance
ssh <gcp-gateway> 'gcloud compute instances stop $GCP_INSTANCE --zone=$GCP_ZONE'Set GCP_PROJECT, GCP_INSTANCE, GCP_ZONE in ~/.config/dsearch/.env.
Scripts: scripts/gpu_embed.py, scripts/gpu_embed_multi.py.
Why two venvs?
GTE requires transformers==4.49 due to a RoPE embedding bug introduced in 4.50. All other models work fine with newer versions. Keeping them separate avoids dependency conflicts.
Why Claude CLI instead of the API?
The answer pipeline uses claude -p (Claude Code headless mode) so it inherits the user's existing subscription. No separate API key or billing setup needed for the LLM steps.
Why HyDE? (optional, --hyde flag)
HyDE was tested as a default and rejected — net -5.7pp (gained 3 abstract queries, lost 7 navigational). It remains available as a manual --hyde flag for abstract/conceptual queries where vocabulary mismatch is the main barrier.
Why agent ranking instead of a reranker?
Cross-encoder rerankers require running a model locally per (query, doc) pair — expensive at 100 documents. Claude agents read the actual extracted text and score on relevance to the intent, not just lexical similarity. Running 10 agents via subscription costs nothing extra per query — no per-token billing — while spinning up a GPU reranker would require local hardware or cloud GPU time.
Why 5-arm RRF instead of a single reranker?
Multi-signal weighted RRF outperforms Cohere Rerank by 12 points on Recall@20. Each signal captures different failure modes: vector catches semantic similarity, TF-IDF catches exact term matches, Recoll path exploits filesystem hierarchy (folder names, dates in paths). No single model captures all three — fusion is strictly better. (FTS5 was tested as a 5th active arm but proved the lexical axis is saturated — disabled.)
Why dual-embed instead of model swap?
The BGE-M3 pilot showed 10/32 miss rescues with 0 Cohere wins, but a regression test revealed 5/73 hits would break on a full swap. Dual-embed (BGE-M3 primary + Cohere optional) captures the rescues while the retained Cohere arm protects current hits when enabled. Trade-off: 2x storage and slightly higher query latency (~1.5s vs ~0.5s).
| Requirement | Version | Notes |
|---|---|---|
| Python | 3.10+ | |
| Recoll | any recent | BM25 index |
| Claude CLI | latest | dsearch-answer only |
| Cohere API key | — | optional (--cohere flag) |
pdftotext |
— | poppler-utils |
antiword |
— | legacy .doc files |
| faiss-cpu | 1.7+ | |
| chromadb | 0.5+ | |
| transformers | ==4.49 (GTE venv) | |
| sentence-transformers | latest | |
| numpy | latest | |
| json-repair | latest | LLM JSON fault tolerance |
| regex | latest | Unicode word boundaries |
MIT