Skip to content

v0.2 corpus model: single typed graph over frozen sources; traceability as a derived view (epic) #157

Description

@luofang34

Epic for the v0.2 corpus rewrite. Milestones v0.2-M1 … v0.2-M7 carry the phase scopes; every phase PR follows the trace-first convention and files its own issue against its milestone.

Problem

The SDLS pilot (onboarding CCSDS 355.0-B-2 / 355.1-B-1 as a downstream adopter) had to hand-roll six schemas Evidence can neither load nor validate — a document registry with hashes, requirement↔normative-text excerpts with page/section locators, extraction provenance + review batches, ambiguities, decisions, and PICS tables. The strain points inside the trace model itself:

  • Lifecycle state smuggled through the unvalidated free-text category field ("candidate:mandatory"); the pilot's own SYS-REVIEW-001 ("model extraction cannot approve requirements") is expressed as data the tool cannot check, and nothing enforces its review batches.
  • Provenance as unqueryable prose (source = "CCSDS-355.0-B-2 section 4.1.1.1.4, PDF page 35, document page 4-2").
  • The same normative sentence duplicated in excerpts and in HLR rationale with no digest binding them — an accidental dual-write.
  • Ingestion via pdftotext -layout dumps, with PNG screenshots as the workaround for destroyed tables (PICS), figures, and conventions pages.

v0.2 makes the corpus the single typed model: frozen sources → committed source graph → reviewed requirements → SYS/HLR/LLR → code/test/result, with traceability as a derived view.

Settled direction

  1. Delete-last convergence. The legacy trace model is generalized into the corpus graph, not deleted first. Evidence dogfoods its own trace; the floors ratchet, trace self-validate, RULES↔LLR bijections, and bundle determinism gates must stay green throughout. The legacy loader is removed only at M6, behind a parity gate.
  2. Lifecycle lands once, on the corpus store. No interim lifecycle on the current model; the SDLS pilot waits for M2.
  3. v0.2 of this repo, in place. No parallel v2 crate; version 0.1.x → 0.2.0 at cutover.
  4. Ambiguities and decisions are first-class graph nodes, in scope for M5.

Design decisions

DD-1 — Single corpus model; traceability is a derived view

One typed, uid-keyed graph is the source of truth. Matrix, floors, bijections, coverage/gap reports, and the context query are computed views over it. End state has no compatibility projection and no dual-write. Rejected: keeping cert/trace as a peer model with sync — that recreates the excerpt/rationale duplication the pilot suffered.

DD-2 — Convergence path: generalize, then cut over (delete-last)

The current four-file trace is already a corpus subset (uid nodes, typed edges). M1 loads it into the graph unchanged; M6 migrates Evidence's own cert/trace into the corpus layout and deletes the legacy loader in the same PR as the parity guardrail. Parity gate: floor values preserved exactly; all bijection lock-tests pass; trace-matrix and golden fixtures byte-identical (or intentionally regenerated with review); check / context outputs equivalent. Rejected: delete-first big-bang — leaves main red for the whole rebuild, zeroes ratchet floors, and removes the place where the rewrite's own trace chain must live.

DD-3 — Store: cert/corpus.toml index over linked TOML files

corpus.toml holds schema_version plus glob lists per node kind (sources, source-graphs, requirements, ambiguities, decisions, profiles, reviews, tests). File layout is non-semantic; the loader unions all files into one graph. Schemas are strict (unknown fields rejected), with per-file [schema] version and the floors policy: an older tool refuses a newer schema rather than skipping fields. Target layout (post-M6):

cert/
  corpus.toml
  sources.lock
  sources/            source-graph/        requirements/{source,sys,hlr,llr}/
  ambiguities/        decisions/           profiles/
  pics/               reviews/             tests/

results/ is intentionally absent: run results are bundle-scoped graph overlays (see DD-15), not committed baseline files.

DD-4 — Node identity and uid scheme

uid is permanent identity; human id is unique per kind and renameable. New corpus node kinds use typed-prefix uids (src_, snode_, patch_, req_, rev_, amb_, dec_, prof_) over a UUIDv4 core, so cross-references self-document and the validator can type-check edges cheaply. Legacy bare-UUID trace uids remain valid until M6, when the migration rewrites them into prefixed form. Locator identity hierarchy for source material: explicit spec ID/numbering → section path + local ordinal → content hash/structural fingerprint → page/DOM/line positions, which are diagnostic only, never identity.

DD-5 — Edges are typed and live on the owning node

derives_from (req→req), quotes (req→source node + span), verifies (test→req), reviews (review→requirement or curated patch @ digest), resolves (decision→ambiguity), concerns (ambiguity→source node | req), supersedes (document revision chains). Edges stay embedded in the owning node's record (as traces_to is today) — one owner per edge, clean diffs. Rejected: a separate edge file (merge-conflict magnet, no owner).

DD-6 — Source freezing (M3)

A baselined source is never just a URL. Registry entry: id, media_type, canonical location, retrieved_at, sha256, capture mode. Capture modes: vendored (raw bytes kept; high-assurance default), hash-only (digest + location; redistribution-restricted documents), external-controlled (immutable ID in an org document system). sources.lock pins the resolved digest set. Content change at the same URL ⇒ new document revision node (supersedes edge); silent update is a validation error. The pilot's "missing transitive normative reference" case is representable as a registry entry with availability = missing — trackable, lintable.

DD-7 — Ingestion contract (M4)

Ingesters are reproducible given (frozen bytes, pinned tool + version) — not deterministic-forever; extractor versions are recorded in the source record. The committed source graph is the reviewed artifact; re-ingestion is a drift lint, not the source of truth. Where a parser fails structurally (PDF tables, notably PICS), a reviewed curated patch layer overrides parser output; patches are first-class records with the same review lifecycle. PDF extraction delegates to a pinned external extractor (pure-Rust PDF text extraction is not adequate for CCSDS layouts); this is an ingest-time-only dependency — downstream consumers only need the committed graph. Ingester order: Markdown → HTML → PDF (cheapest first proves the node/locator design before the hardest format).

DD-8 — Source graph schema

SourceNode { uid, source_revision, parent, kind, ordinal, label, text, content_sha256, locator } with kinds: Section, Paragraph, ListItem, DefinitionTerm, DefinitionBody, Table, TableRow, TableCell, CodeBlock, Note, FigureCaption. Locators per format (PDF page/section/paragraph/bbox; HTML canonical_url/fragment/heading_path/dom_path; Markdown path/git_blob/anchor/heading_path/byte_range). Prose normalization at ingest is Unicode NFC + whitespace-run folding + trim; code nodes preserve significant spaces and line boundaries while normalizing Unicode and line endings; no automatic dehyphenation (too risky for normative text — a curated patch fixes real cases). content_sha256 is over normalized text; requirement quote spans index into normalized node text, so quote digests are stable by construction.

DD-9 — HTML and Markdown are structure-preserving

HTML ingestion keeps the h1–h6 tree, <p>/nested lists, <dl>/<dt>/<dd>, table row/column structure, <code>/<pre> literals, ids/anchors/internal links, and note/example/figure classification. Scripts, styles, and inert templates are excluded by closed rule; navigation, ToC duplicates, headers, and footers require committed recipe selectors so potentially normative content is never dropped heuristically. Markdown via CommonMark/GFM AST (headings, lists, tables, blockquotes/admonitions, fenced code, footnotes, explicit heading IDs); local files pin the git blob SHA, remote ones pin content SHA-256 + final URL + retrieval metadata. Rationale is the OIDC Core case: RFC 2119 notation section, monospace-means-literal, and Terminology definitions that are normative without any capitalized keyword — unreachable from a text dump or a MUST-regex.

DD-10 — Requirement records (M5)

layer = source | sys | hlr | llr; one requirement may cite multiple [[sources]] bindings, each document + node + span + quoted_text_sha256; a single source node may yield many atomic requirements, but each must point at its precise node/span, never "the chapter". Canonical modality enum: required | prohibited | recommended | not_recommended | optional | permitted | normative_definition | informative; per-document conventions map the spec's vocabulary (RFC 2119, CCSDS shall/should/may/permits) onto it. [semantics] behavior tags are optional and non-gating initially (an unvalidated vocabulary must not become a load-bearing surface by accident). [extraction] metadata (agent, prompt digest, tool version) is mandatory on machine-proposed candidates.

DD-11 — Lifecycle and review (M2)

States: candidate → approved | rejected, plus stale (derived). Approval is a separate review record binding (requirement uid, reviewed_content_sha256, decision, reviewer, reviewed_at); any content change flips the requirement to stale automatically because the digest no longer matches. Agents may create and update candidates only, via an append-only proposal path; approval, source-snapshot mutation, and baseline overwrite are human-only. Enforcement: in strict profiles, implementation artifacts (LLR modules, code, tests) may only trace into approved requirements; a candidate:-style prefix in category becomes a lint error after M6.

DD-12 — Conventions baseline precedes extraction (M5)

Before body extraction, per-document records must be extracted and approved: requirements_notation (vocabulary, case sensitivity) and a normativity map (default classification, informative sections, per-node overrides — e.g., OIDC Terminology marked normative with a reason). Processing order: freeze → source graph → conventions review → normativity review → batched candidate extraction → atomicity/completeness/conflict lints → human approval → SYS/HLR/LLR derivation. The modal-verb scan is a completeness lint (flags normative-looking nodes not covered by any requirement), never the extraction algorithm.

DD-13 — Ambiguities and decisions as first-class nodes (M5)

Ambiguity: id, status, severity, concerns edges, question, required_resolution. Decision: id, status, statement, rationale, source, optional resolves edges. Validation: an open blocking ambiguity prevents approval of requirements whose cited nodes it concerns; release-grade profiles gate on zero open high-severity ambiguities. Decisions get the same digest-bound review treatment as requirements.

DD-14 — Profiles and PICS

The applicability/profile filter ships in M5 (the pilot's first milestone is profile-scoped — "TC AEAD baseline" — and honest gap reports need it); full PICS form modeling/rendering ships in M7. PICS items link to the requirements they claim; claim_status cannot be asserted while linked requirements are unapproved or ambiguity-blocked.

DD-15 — Derived views, floors, bundles

All reports (matrix, coverage, gap, context) are graph queries. Floors gains corpus dimensions post-cutover (e.g., frozen source count, approved source-requirements per document; open high-severity ambiguities as a ceiling for release profiles) while trace_* dimensions carry over with values preserved exactly. Corpus artifacts (graph files, requirements, reviews, sources.lock) enter the bundle's SHA256SUMS/verify surface; vendored source bytes are referenced by hash from sources.lock, never copied into bundles. Run results remain bundle-scoped overlays, not committed corpus records — committed baseline stays pure intent, bundles stay the evidence of execution.

DD-16 — Validation is phase-aware and incremental

A corpus declares its phase per document/workstream (e.g., frozen, graphed, conventions-approved, extracting, baselined); validators key on the declared phase so "not yet extracted" is distinguishable from "invalid" — and phase claims are themselves checked (a baselined document with zero approved requirements is an error). Per-document content-hash short-circuiting keeps corpus validate fast on large graphs (a 129-page PDF or OIDC-sized HTML yields thousands of nodes).

DD-17 — CLI and MCP surfaces (M7, verbs earlier as needed)

CLI verb family: cargo evidence corpus source add|inspect, corpus ingest, corpus validate, corpus queue, corpus context, corpus review, corpus render. MCP: read-only evidence_corpus_validate | _source_context | _extraction_queue | _requirement_context | _trace, plus the restricted candidate-proposal API. Every new verb registers in KNOWN_SURFACES with a matching HLR (surface bijection), and JSONL invariants (single terminal, stdout-strict) apply unchanged.

DD-18 — Code placement and repo discipline

Corpus code lives in evidence-core under a corpus.rs + corpus/ module tree (domain-named submodules; 500-line file limit; no mod.rs). Existing conventions bind: trace-first chain seeding per PR, one issue per PR, fix+guardrail in the same PR, walkdir rules, thiserror-typed errors, no panics in library code. The design spec lands as docs/superpowers/specs/2026-07-18-corpus-model-v0.2-design.md in the M1 seed PR.

Milestones

Milestone Scope anchor
v0.2-M1 Corpus graph core DD-1..5, DD-18
v0.2-M2 Lifecycle & review DD-11
v0.2-M3 Source layer DD-6
v0.2-M4 Ingesters DD-7..9
v0.2-M5 Extraction & corpus semantics DD-10, DD-12..14
v0.2-M6 Cutover DD-2 parity gate; legacy loader deleted
v0.2-M7 Surfaces DD-17, PICS forms

Acceptance fixtures

  • OIDC Core HTML: Requirements Notation section; a Terminology definition normative without capitalized keywords; ID Token claims (one node → several atomic requirements with distinct spans).
  • An equivalent small Markdown spec (heading IDs, tables, fenced code).
  • An SDLS PDF fragment with document page labels and paragraph numbering.
  • A CCSDS PICS table — the known parser-hostile case; exercises the curated-patch path end to end.

These prove structured-standard support rather than PDF text scraping.

Non-goals for v0.2

  • No requirement NLP/semantic dedup beyond the declared lints.
  • No live-fetch of sources at validate time (frozen bytes or recorded digests only).
  • No agent-side approval authority anywhere.

Tracking hierarchy

#157 is the v0.2 parent epic. Native subissues provide the milestone roadmap:

Only the active milestone receives implementation child issues. Every implementation PR has one dedicated issue and includes its regression guardrail. M1 through M3 are complete.

Active M4 tracker #167 owns implementation issues #216 through #222 in contract-first landing order: source-graph schema, Markdown, HTML, curated patches, review-target generalization, drift, then PDF.

Assurance-correctness foundation

The v0.1.6 Assurance correctness milestone is complete. Issues #139 and #141 through #145 establish recipe/output identity, fail-closed assurance semantics, schema-valid initialization, locked/offline execution, complete assurance diff/completeness, and explicit assurance selection.

These contracts are inputs to M4 through M7 and remain protected by their regression gates.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions