Multi-source knowledge base RAG platform with automatic incremental sync.
Builds on production-rag-platform's
proven core pipeline (chunking, hybrid dense+sparse search, cross-encoder
reranking, citation-aware generation, OpenTelemetry tracing) rather
than rewriting it from scratch. The new value-add here: ingesting multiple
document types (PDF, Markdown, Notion) through a shared Connector
interface — plus a standalone web-page parser (app/parsing/web_parser.py)
not yet wired into a connector, see Known Limitations
— and automatic incremental re-sync: unchanged, healthy content is
skipped, and detected index drift (points missing, partially missing,
or orphaned in Qdrant) is automatically repaired rather than silently
trusted.
Full sprint-by-sprint plan: docs/PLANNING.md.
- Multi-source ingestion — filesystem (PDF/Markdown) and Notion behind
one
Connectorinterface, with hash-based incremental sync (skip unchanged, update changed, delete vanished — no orphan chunks) - Hybrid search — dense + sparse (BM25) retrieval with native Qdrant RRF fusion, cross-encoder reranking
- Source-scoped citation validation — every citation is checked against
the exact
(source_type, source_id, location)it was generated from, so two different documents can never spoof each other's citations (proven with a dedicated cross-source-leak test, not just claimed) — this is citation integrity, not semantic grounding; see Citation integrity validation, not semantic grounding - Full distributed tracing — a sync run or a chat request is one Jaeger trace end to end, cross-checked span-by-span against Jaeger's own API, not just visually inspected
- LLM-judged evaluation harness — DeepEval + a local
qwen2.5:7bjudge, reporting retrieval/generation metrics broken down by content type from a real golden set - One-command deployment —
docker compose upbrings up the full stack; sync history survives a container restart (a real named volume, verified, not assumed) - 340+ tests, almost all against real dependencies (real SQLite files, a real Qdrant instance, real Jaeger, real browser automation for the UI) rather than mocks — see Known Limitations for what's honestly not covered yet
- Status
- Architecture
- Beyond the Happy Path
- Technologies Used
- Quick start (Docker Compose)
- LLM providers
- Connectors
- Sync
- Citation format
- UI
- Development setup
- Known Limitations
- License
Sprints 0–17 complete, including 17.1–17.4 hardening — the platform is fully working end to end:
- Core RAG pipeline (parsing, hybrid search, reranking, citation-aware generation) ported from production-rag-platform
- Multi-source ingestion: filesystem (PDF/Markdown) + Notion connectors
behind a shared
Connectorinterface - Incremental sync (skip/update/delete by content hash) with a scheduler, manual trigger, and full history — including a zero-downtime versioned re-index with automatic rollback of a partially-written new version on mid-batch failure (see Sync)
- Distributed tracing (OpenTelemetry + Jaeger) across both sync and chat
- Golden-set evaluation (DeepEval + a local judge model)
- A multi-page Streamlit UI (Chat, Sources, Sync Status)
- One-command Docker Compose deployment, with sync data surviving a container restart
- CI (lint + real-Qdrant test suite + Docker build) on every push
See docs/PLANNING.md for the sprint-by-sprint closing notes each of these is backed by, and Known Limitations below for what's not covered — most notably, the Notion connector has never been tested against a real workspace (no API key available).
graph TD
subgraph Host["Host machine"]
Ollama["Native Ollama<br/>generation + embedding"]
UI["Streamlit UI<br/>.venv-ui, make ui"]
DocsFolder["./data/documents<br/>bind-mounted"]
end
subgraph Compose["docker compose"]
Backend["FastAPI backend<br/>app.server:app"]
Scheduler["SyncScheduler<br/>periodic, per-connector interval"]
Manager["SyncManager"]
Connectors["Connectors<br/>filesystem, notion"]
Registry[("registry.db<br/>documents + sync_runs<br/>named volume")]
Qdrant[("Qdrant<br/>hybrid dense+sparse index")]
Jaeger["Jaeger<br/>OTLP traces"]
end
UI -->|"POST /chat, /sync/*, GET /sources"| Backend
Backend -->|"embed + generate"| Ollama
Backend -->|"hybrid search"| Qdrant
Backend -->|"OTLP spans"| Jaeger
Backend -. "starts on boot - lifespan" .-> Scheduler
Scheduler --> Manager
Manager --> Connectors
Manager --> Registry
Manager --> Qdrant
Connectors --> DocsFolder
This section exists because "ingest a PDF, put it in Qdrant, ask an LLM" is the easy 80%. Seven rounds of independent code review (Sprints 12, 16, 17, 17.1, 17.2, 17.3, 17.4) found real bugs in the other 20% — the kind that only show up once a system has to survive its own edge cases, not just its demo path. Each item below is a real bug, found and fixed, not a hypothetical.
- Point identity, not just point storage.
QdrantStore.point_id_for's key was built fromdoc_id(a content hash) withoutsource_id— two different documents with byte-identical content (e.g. a PDF duplicated under two filenames) silently collided on the same point ID, and the second upsert overwrote the first with no error. The review that found this also found the existing regression test was itself hiding the bug — a shared default in a test helper meant both "different" documents in the test were already colliding before the test's own assertions ever ran. Both the bug and the test that failed to catch it were fixed together (Sprint 17), with a new end-to-end test proving two identical-content files keep independent registry and Qdrant identity. - Versioned re-index with real cancellation safety. A re-index
embeds and upserts a document's new version before deleting the
old one, so a failure mid-embed leaves the old version searchable
instead of the document going dark (Sprint 13). A later review found
that guarantee only held within a single batch — a multi-batch
failure could leave a partial new version stranded forever — and
that
asyncio.CancelledError(a real app-shutdown or scheduler-stop signal) bypassed the rollback entirely, since it inherits fromBaseException, notException(Sprint 16, Sprint 17). Both are proven with a realtask.cancel()delivered mid-embed viaasyncio.Event, not a manually-raised substitute standing in for one. - Registry and Qdrant are two separate stores that can drift apart —
and now self-heal. Incremental sync originally trusted a single
signal: "has the content hash changed?" A review pointed out that
Qdrant's data can disappear by means the app never sees (manual
deletion, external tooling, partial data loss) while the registry's
hash stays exactly the same, so a document could silently stay
unsearchable forever.
QdrantStore.has_document_version()andcount_for_document_version()now reconcile the two on every sync (Sprint 17.2) — proven by deleting a document's Qdrant points directly, leaving the registry untouched, and watching the very next sync detect and repair it automatically, with no manual intervention. - A schema migration tested against the schema it actually has to
migrate, not a simplified stand-in. SQLite has no
ALTER COLUMNto relax aNOT NULLconstraint. A migration meant to make achunk_countcolumn nullable was validated against a fixture that simulated "the column doesn't exist yet" — but a real database from the previous sprint already had the column, asNOT NULL DEFAULT 0, so the migration was a silent no-op against it and could raisesqlite3.IntegrityErroron a real upgrade. The fix was tested against a fixture that reproduces the actual prior schema byte-for-byte, confirmed to fail first, then rebuilds the table (SQLite's own workaround for the missingALTER COLUMN) to genuinely drop the constraint (Sprint 17.4). - A recurring theme, not a one-off: green tests that hid real bugs. This happened at least three separate times across the hardening sprints — the point-identity test above (Sprint 17), two Qdrant schema-validation tests that started passing for the wrong reason once a new check was added earlier in the same function, masking the dense-vector check they claimed to exercise (Sprint 17.1), and the migration test that only ever covered the easy case (Sprint 17.4). Each time, the fix wasn't just the code — it was proving the test would have failed before the fix, and rewriting the test alongside the bug.
- In real numbers: 426 tests (most against real dependencies — a
real SQLite file, a real Qdrant instance, real Jaeger traces, real
browser automation — not mocks), across seven independent review
rounds. The embedding-concurrency default wasn't guessed: a real
benchmark against native Ollama, later hardened with warmup runs,
repeated samples, and randomized ordering after a review questioned
the first pass's methodology, found
concurrency=4(57.1 mean chunks/sec) andconcurrency=8(55.6 mean chunks/sec) statistically indistinguishable — well within each other's measured variance — so 4 stayed the default rather than doubling open connections for no measured gain.
Layer by layer, what's actually running (not aspirational — each line was verified in a sprint closing note in docs/PLANNING.md):
| Layer | Technology | Notes |
|---|---|---|
| Parsing | PyMuPDF (fitz) — PDF; hand-written heading-block parser — Markdown; trafilatura — web pages |
Page/paragraph extraction ported from production-rag-platform (Sprint 0); Markdown's heading-path location scheme (Sprint 3) is reused by the web parser too (Sprint 6) |
| Chunking | Whitespace token counter, 500/50 (size/overlap) | Provisional default, unchanged since Sprint 0 — never re-tuned against a larger corpus (see Known Limitations) |
| Document registry | SQLite (stdlib sqlite3), no ORM |
(source_type, source_id) primary key, content-hash diffing for incremental sync (Sprint 2) |
| Connectors | LocalFilesystemConnector (PDF/Markdown); NotionConnector (Notion API, 429 retry/backoff) |
Shared async Connector Protocol (Sprint 3, generalized to async in Sprint 6) |
| Sync | Hand-rolled asyncio loop (SyncScheduler) + per-connector concurrency guard |
APScheduler/Celery deliberately rejected — no cron expressions or job persistence needed (Sprint 7); re-index is zero-downtime with deferred cleanup, not atomic (Sprint 13); embedding calls run with bounded concurrency, default 4, picked from a real benchmark that found a throughput plateau past that point — see Sync (Sprint 14) |
| Provider abstraction | ChatProvider/EmbeddingProvider Protocols — Ollama (native) + Claude (Anthropic API) |
Claude has no embedding endpoint, so embedding always stays on Ollama regardless of chat provider (Sprint 1) |
| Embedding | Ollama, nomic-embed-text |
768-dim, cosine distance; search_document:/search_query: task prefixes required for quality |
| Generation | Ollama qwen2.5:7b-instruct (default) or Claude |
Model is a config value, not hardcoded |
| Vector DB | Qdrant | Dense + sparse (BM25 via FastEmbed Qdrant/bm25) hybrid search with native RRF fusion |
| Reranking | sentence-transformers CrossEncoder, ms-marco-MiniLM-L-6-v2 |
Candidate k=20 → top n=5 (Sprint 5) |
| Backend | FastAPI | SSE streaming for /chat; containerized since Sprint 11 |
| Citations | [s.source_type:source_id/location] |
location is page/paragraph for PDF or a heading path for Markdown/web/Notion — checked against the full triple so two sources can't spoof each other's citations; citation integrity, not semantic grounding (Sprint 0/3/5/12, see Citation format) |
| Observability | OpenTelemetry + Jaeger | A full sync run and a full chat request are each a single trace end to end (Sprint 8) |
| Evaluation | DeepEval + qwen2.5:7b-instruct (judge) |
RAGAS was tried and rejected in production-rag-platform for a real dependency conflict — not re-attempted here (Sprint 9) |
| UI | Streamlit, multi-page (st.navigation) |
Separate venv (.venv-ui) — a real, confirmed starlette version conflict with FastAPI's pin (Sprint 10) |
| Orchestration | Docker Compose | Qdrant + Jaeger + backend containerized; Ollama stays native — no Metal GPU passthrough on Docker Desktop macOS (Sprint 11) |
Requires a native Ollama install — Docker Desktop on
macOS has no Metal GPU passthrough, so Ollama runs on the host and the
backend container reaches it via host.docker.internal, not in a
container.
ollama pull qwen2.5:7b-instruct
ollama pull nomic-embed-text
docker compose up -d --build # Qdrant + Jaeger + backendNOTION_API_KEY is optional — set it in a repo-root .env (read by
Docker Compose for variable substitution) before up only if the Notion
connector should actually run; leave it unset and filesystem is the only
active connector, same as running on bare Python.
Drop real files into ./data/documents/ (bind-mounted into the
container), then either wait for the next scheduled sync or trigger one
now:
curl -X POST http://localhost:8000/sync/filesystem
curl http://localhost:8000/health/ollama # proves the container can reach native OllamaThe registry/sync-history SQLite file lives in a named volume
(registry_data) — it survives docker compose restart/stop+start;
only docker compose down -v wipes it (genuinely fresh install).
Generation (chat) and embedding are independent, config-driven choices — Claude has no embedding endpoint, so embedding always stays on Ollama regardless of which chat provider is selected:
GENERATION_PROVIDER=ollama # or claude — local-first default
EMBEDDING_PROVIDER=ollama # only option today
CLAUDE_API_KEY=sk-ant-... # required if GENERATION_PROVIDER=claudeTwo Connector implementations exist: LocalFilesystemConnector
(PDF/Markdown from a local folder) and NotionConnector (real Notion API
calls — search + block-children endpoints, with 429 retry/backoff; needs
NOTION_API_KEY). Both go through the same ingest_connector() incremental
sync pipeline (skip/update/delete), unmodified.
Each connector syncs on its own interval (FILESYSTEM_SYNC_INTERVAL_SECONDS,
NOTION_SYNC_INTERVAL_SECONDS), plus a manual trigger:
curl -X POST http://localhost:8000/sync/filesystem
curl http://localhost:8000/sync/filesystem/historyTwo syncs of the same connector never run concurrently — a second
attempt while one is in progress is rejected immediately (409), not
queued. Every run (scheduled or manual) is recorded with its outcome,
duration, and how many documents changed/were skipped/deleted.
Deliberately not called "atomic" — a true atomic swap would need transactional guarantees Qdrant doesn't offer, and calling it that would be a claim this system can't back up. What it actually does, and why:
When a document's content changes, the new version is fully parsed,
embedded, and upserted into Qdrant first — in batches of
upsert_batch_size — tagged with a document_version payload field
(the new content hash). Only once every batch of the new version is
confirmed upserted does the old version's chunks get deleted. Before
Sprint 13, this was reversed — old chunks were deleted before
re-ingesting — so a failure partway through embedding (a network blip,
an Ollama timeout) left the document with no searchable chunks at all
until a later sync succeeded. Deferred cleanup closes that specific
window: a failed re-index now leaves the old version fully intact and
searchable, never a half-written document.
Sprint 13's fix had its own gap, closed in Sprint 16: the "new version
either fully lands or none of it does" guarantee only held within one
batch. With a multi-batch document, an earlier batch could already be
committed under the new document_version when a later batch's embed
call failed — the exception propagated out of ingest_connector before
cleanup ever ran, leaving a partial new version sitting in Qdrant
forever alongside the (correctly intact) old version. The fix: on any
failure mid-loop, every point already upserted under that
document_version is explicitly rolled back
(QdrantStore.delete_version()) before the exception is re-raised,
restoring "only the old version, nothing from the new one" for the
whole document — proven with a real 3-batch failure scenario, not just
the single-batch case (tests/test_versioned_reindex.py::test_multi_batch_partial_failure_rolls_back_the_partial_new_version).
The honest tradeoff this doesn't eliminate: between the new version's upsert finishing and the old version's cleanup running, both versions are simultaneously present and searchable — a query in that window can return duplicate/stale-alongside-fresh chunks for the same document. This window is real, not hypothetical (proven with a test that inspects Qdrant's actual contents at that exact moment), and its duration scales with how many batches a re-index takes, not a fixed number: measured at ~12–20 microseconds for a single-batch document (bounded by the time between two sequential Qdrant calls — all embedding happens before the window opens), but ~1.5–3 milliseconds across several runs for a real 7-batch document — because the window actually starts at the first batch's upsert, not the last, so a multi-batch re-index's new chunks are partially visible for the whole remaining ingestion time, not just a last-call gap. See the Sprint 13 and Sprint 16 closing notes in docs/PLANNING.md for both measurements and the before/after failure-scenario proof.
Embedding calls during ingestion run with bounded concurrency
(EMBEDDING_CONCURRENCY, asyncio.Semaphore) instead of one at a time.
The real question — does a single native Ollama instance actually get
faster with more concurrent requests, or does it queue/degrade past some
point? — was benchmarked directly
(scripts/benchmark_embedding_concurrency.py, nomic-embed-text on an
M2), not assumed. Sprint 14's original run took one sample per
(chunk_count, concurrency) pair; an external review correctly flagged
that "plateau within measurement noise" wasn't backed by an actual
variance number. Sprint 16 hardened the methodology — a warmup call per
chunk count, 3 repeats per pair with concurrency order randomized each
repeat, mean/median/stddev reported — and re-ran it for real:
| Concurrency | mean chunks/sec | median | stddev | n |
|---|---|---|---|---|
| 1 | 26.9 | 28.2 | 7.6 | 9 |
| 2 | 48.9 | 58.0 | 19.8 | 9 |
| 4 | 57.1 | 63.6 | 18.5 | 9 |
| 8 | 55.6 | 62.2 | 23.8 | 9 |
(n = 3 repeats × 3 chunk counts of 10/100/1000, each repeat's concurrency order shuffled independently.)
The result is a plateau, not unbounded scaling — exactly the failure
mode a single-model native Ollama instance could plausibly hit, so it was
worth actually measuring rather than assuming "more concurrency = more
throughput," and this time the plateau claim is backed by real variance:
1→2 is a genuine jump (48.9 vs 26.9, a gap far larger than either's
stddev); 2→4 is a smaller further gain; 4→8 is not distinguishable
from noise — 57.1 vs 55.6 mean, well inside both configurations'
stddev (18.5 and 23.8) — while holding twice as many connections open
for it. EMBEDDING_CONCURRENCY defaults to 4 — same choice as
Sprint 14, now confirmed by a statistically honest re-run rather than a
single sample per point.
A real sync run's own time breakdown (7 chunks, EMBEDDING_CONCURRENCY=4,
captured from the same OTel spans Jaeger uses — Sprint 8):
| Stage | Duration |
|---|---|
| Total sync | 939 ms |
| Embedding (Ollama) | 812 ms (86%) |
| Qdrant upsert | 29 ms (3%) |
| Parse + chunk | 2 ms (<1%) |
Embedding dominates, as expected — Qdrant's own write path is fast and not the bottleneck worth optimizing further.
Citations are multi-source from the start:
[s.<source_type>:<source_id>/<location>]
examples: [s.filesystem:handbook_pdf/2/0] (PDF: page/paragraph)
[s.filesystem:readme_md/Kurulum/Adım 1] (Markdown: heading path)
source_type identifies the connector a document came from (filesystem
today; notion/confluence later), not its file format — the same
connector can ingest multiple formats. Every citation is checked against
the full (source_type, source_id, location) triple, so two different
sources can safely share the same location without one masquerading as
the other.
app/llm/grounding.py::check_grounding proves every citation tag in an
answer points to a chunk that was actually in the retrieved context, from
the source it claims. It does not prove the specific claim next to
that citation is actually supported by the chunk's text — a model could
cite a real, correctly-attributed chunk beside a claim that chunk doesn't
support, and this check still reports it as grounded. This is citation
integrity validation, not semantic grounding, and the UI/API name it
accordingly (GroundingResult.grounded requires both has_citations and
citations_valid — an answer with zero citations is not grounded, the
most dangerous hallucination shape since there's no citation tag at all
to question).
Future work: claim-level semantic support checking (e.g. an
NLI/entailment check between each claim and its cited chunk's text) would
close this gap — not attempted yet. See the Sprint 12 closing note in
docs/PLANNING.md for the real bug this sprint fixed
(a citation-free answer used to be reported grounded: True).
A multi-page Streamlit UI (Chat, Sources, Sync Status) runs in its own
venv (.venv-ui) — streamlit==1.61.1 and this project's fastapi==0.115.6
pin have a real, unresolvable starlette version conflict (confirmed via
pip install, see the Sprint 10 closing note), so the UI can't share the
backend's venv:
python3.12 -m venv .venv-ui
.venv-ui/bin/pip install -r requirements-ui.txt
make ui # streamlit — points at BACKEND_URL (default http://localhost:8000)It works against either backend: the containerized one (docker compose up,
above) or a host-run one (make dev, below) — it's a pure HTTP client,
never importing backend code directly, only POST /chat, GET/POST /sync/..., GET /sources, and Jaeger's own HTTP API.
Requires Python 3.11+ and a native Ollama install (same host-only reasoning as the Docker Compose setup above).
python3.12 -m venv .venv
.venv/bin/pip install -r requirements-dev.txt
docker compose up -d qdrant jaeger # just the two stateless services
ollama pull qwen2.5:7b-instruct
ollama pull nomic-embed-text
make dev # real backend on the host: uvicorn app.server:app --reloadRun the test suite:
make testTests that require live Ollama/Qdrant skip automatically when those services aren't reachable.
Real, documented gaps — not a hedge. Each one is traceable to a sprint closing note in docs/PLANNING.md, or to a direct check of the code/tests where noted.
-
Re-indexing a changed document is zero-downtime with deferred cleanup, not atomic — during re-index there's a window (measured at ~12–20 microseconds locally for a single-batch document, but ~1.5–3 milliseconds for a real multi-batch document, since the window opens at the first upsert_batch, not the last — see Sync) where both the old and new version of a document's chunks are simultaneously searchable, so a query in that window can see duplicate/stale-alongside-fresh results. This is a deliberate, measured tradeoff (Sprint 13) in exchange for closing a worse problem (Sprint 4's original ordering could leave a document with zero searchable chunks if a re-index failed mid-embed) — see the Sync section above. A related gap (Sprint 17): a re-index cancelled mid-batch, or hitting a failure that spans more than one batch, now correctly rolls back the partial new version regardless — see the Sync section's rollback description.
-
The Notion connector is mock-tested only, and doesn't recurse into nested blocks. No
NOTION_API_KEYwas available on the development machine (Sprint 6), so its 16 tests simulate Notion's real documented JSON shapes (search pagination, block-children pagination, 429+Retry-After, other errors) viahttpx.MockTransport— a real end-to-end run (tests/test_notion_e2e.py) exists and will run automatically once a key is set, but hasn't yet. Separately, and regardless of that: block extraction (app/connectors/notion.py) only reads text from a fixed set of top-level block types (paragraphs, list items, quotes, to-dos, code, headings) and is deliberately non-recursive into nested children — a page with content inside a toggle, a nested list, or a synced block loses that content entirely, the same restraintLocalFilesystemConnectorapplies to nested folders. -
The Claude API path has never been compared against Ollama on a real question. No
ANTHROPIC_API_KEYwas available (Sprint 1) —tests/test_provider_comparison_e2e.pyauto-skips. This is "the test never ran," not "no difference was found" — the comparison stays an open question. -
Cross-lingual query/content retrieval is measurably weaker, and the reranker makes it WORSE — confirmed with an isolated experiment, not just a hypothesis. Sprint 17.5 first noticed the eval CLI (
app.evaluation.cli) used to measure hybrid retrieval BEFORE reranking — a different pipeline than a real chat query goes through (app/wiring.pyalways passes aCrossEncoderRerankertosearch()) — wired the same reranker into the CLI (default ON, matching production;--no-rerankeropt-out for the old pre-rerank measurement), and observed PDF recall drop from 0.429 to 0.143 with reranking on, floating "Turkish golden set + English-trained reranker" as an unconfirmed guess at the cause. Sprint 17.7 checked that guess directly against the fixtures rather than assuming it: the golden set isn't uniformly Turkish — all 12 questions are Turkish, but the PDF source document is entirely English and the Markdown source document is entirely Turkish, so the PDF half was already cross-lingual (Turkish question, English content) while the Markdown half was mono-lingual (Turkish question, Turkish content) — and only the cross-lingual half had regressed. That reframed the guess into a falsifiable prediction (mismatch drives the regression, not Turkish specifically) and Sprint 17.7 tested it directly: a parallel English question set (tests/fixtures/golden_set_en.json, direct translations, identicalexpected_locations— content unchanged) turns the PDF half mono-lingual and the Markdown half cross-lingual, giving a full 2×2 (content × question language) × reranker-on/off design, 8 real cells against native Ollama + Qdrant:no rerank reranked pairing PDF + Turkish question recall 0.429 recall 0.143 cross-lingual PDF + English question recall 0.857 recall 0.857 mono-lingual Markdown + Turkish question recall 1.000 recall 1.000 mono-lingual Markdown + English question recall 1.000 recall 0.800 cross-lingual Both cross-lingual cells dropped under reranking; both mono-lingual cells were completely unchanged (precision too, to the decimal) — a clean, repeated pattern across all four cells, not a coincidence in one. A second, independent signal: PDF's mono-lingual pre-rerank recall (0.857) is already dramatically higher than its cross-lingual pre-rerank recall (0.429) —
nomic-embed-textitself retrieves worse across languages before the reranker (cross-encoder/ms-marco-MiniLM- L-6-v2, English-trained) ever runs, so reranking sharpens an existing weakness rather than creating a new one. Caveat, stated as plainly as the finding itself: this is one golden set (12 questions, 2 documents, one fictional domain), one run, no statistical significance testing — a real, reproduced finding for this specific reranker/embedding-model/ golden-set combination, not a general claim about cross-lingual rerank performance everywhere. See the Sprint 17.7 closing note indocs/PLANNING.mdfor the full 8-cell breakdown. -
The root cause of PDF's weaker retrieval within mono-lingual pairs (PDF+English recall 0.857 vs. Markdown+Turkish recall 1.000 — page-level chunk granularity vs. Markdown's heading-scoped blocks giving the retriever a harder or easier target?) was observed in Sprint 9 but not investigated further.
-
No
WebConnectorexists — only the web page parser (app/parsing/web_parser.py,trafilatura) and its chunker are built and tested against a real HTML fixture (Sprint 6). There's no ingest/discovery path from a URL list into the registry/sync pipeline; that sprint's DoD was parsing only. -
No Confluence connector. Sprint 18 (a second connector, proving the
Connectorabstraction generalizes) is a stretch goal and hasn't been attempted yet. -
The sync concurrency lock is process-local, not distributed —
SyncManager._running(Sprint 7) is a plaindict[str, bool]on one Python object, so it only prevents overlapping syncs for the same source within a single process. Running multiple worker processes or replicas against the same registry/Qdrant would let two of them sync the same source concurrently with no cross-process coordination (no distributed lock, e.g. via Postgres/Redis). Currently latent, not manifesting: the Dockerfile'sCMDruns a singleuvicornprocess with no--workersflag. -
A cosmetic tracing gap: every sync/ingestion span's
otel.scope.nameshowsapp.sync.managerin Jaeger, regardless of which module actually produced it (most are reallyapp.ingestion.ingest) —SyncManagerpasses its own tracer down explicitly so tests can capture spans with an isolatedTracerProvider. Doesn't affect span hierarchy, names, or attributes, just that one instrumentation-scope label (Sprint 8). -
No authentication, authorization, or multi-tenancy anywhere in the API — verified directly against the code, not just absent from a sprint's scope: no auth middleware or per-tenant scoping exists in any
app/api/*.pyroute. Every endpoint (/chat,/sync/*,/sources,/health*) is open to anyone who can reach the port. -
POST /sync/{source_type}is synchronous — it blocks until the whole sync finishes rather than returning a background-job id immediately. A deliberate choice (Sprint 7): a real need for fire-and-forget syncing over many/large documents hadn't shown up yet. -
Golden-set retrieval precision has a structural ceiling that can be misread as a quality problem:
search()returns the top 5 chunks by default (RERANK_TOP_N), and each Sprint 9 golden question has exactly one expected location — so even perfect retrieval caps precision at 1/5 = 0.2. Markdown's questions actually hit that ceiling every time; read the Sprint 9 closing note before comparing precision numbers across differenttop_nconfigurations. -
Chunk size (500/50 tokens) and rerank k/n (20/5) are untuned defaults, carried over from Sprint 0 and never revalidated against a larger or more diverse corpus than the golden set's two small fixture documents.
-
Single-session UI, no persisted conversation history —
st.session_stateholds Chat page history only for the current browser session; a refresh clears it (consistent with the no-multi-tenancy point above). -
list_source_ids()'s per-sync cost is O(total chunks), not O(document count) — the Qdrant-only orphan cleanup added in Sprint 17.3 scans every point's payload for a givensource_type(a paginatedscroll), so a source with many documents and many chunks per document pays for a scan proportional to its total point count on every sync, not just its document count. Disclosed, not measured against real Qdrant at scale this project has never reached. -
Payload indexes (Sprint 17.3) aren't applied retroactively —
ensure_collection()only creates thesource_type/source_id/document_versionkeyword indexes for a brand-new collection; an existing collection created before Sprint 17.3 shipped (or before an upgrade) keeps running without them, since an existing collection that already passes schema validation is never mutated — consistent withUnexpectedCollectionSchemaError's "don't touch an existing collection" policy elsewhere in this file.
MIT — see LICENSE.