Skip to content

Read document structure from what a document states, not only from how it looks - #210

Merged
MikeAlhayek merged 9 commits into
mainfrom
general-document-structure-engine
Sep 21, 2026
Merged

MikeAlhayek merged 9 commits into
mainfrom
general-document-structure-engine

Conversation

@MikeAlhayek

@MikeAlhayek MikeAlhayek commented Sep 20, 2026 •

Copy link
Copy Markdown
Member

Why

TocSeededStructureAnalyzer was a magazine analyzer with a safe fallback. It read a contents page, else inferred headings from type size, else gave up and called the file one article. Three properties stopped it generalising:

  • One boundary per page. The topmost headline opened an article and every other story on that page was absorbed into it — so a newspaper page of five stories became one.
  • No hierarchy. DocumentStructure.Articles was flat and there was no section type, so a user manual's part → chapter → procedure collapsed into one article.
  • Type size was the only structural signal. Only the PDF reader wrote anything an analyzer could use. Word and HTML wrote nothing at all, so every .docx and every crawled page ingested as one undivided article — despite both stating their structure far more plainly than any PDF does.

What changed

IDocumentStructureAnalyzer now runs a ladder, each rung authoritative over the next, each degrading to the one below. The contract that nothing may fail an ingest is kept: the last rung is still one division covering everything.

Rung DocumentStructureSources Authority Covers
1 Outline stated manuals, reports, books, most white papers
2 StatedHeadings stated Word, HTML, tagged PDFs, Document Intelligence
3 TableOfContents stated, fuzzy-matched magazines, journals
4 InferredHeadings inferred anything with type-size contrast
5 Whole none the answer that is never wrong

Two authoritative sources shipped inside PdfPig and nothing here touched them. TryGetBookmarks gives a hierarchical outline with page destinations — the document stating its own structure, which no heuristic can beat. Page.GetMarkedContents gives real roles on tagged PDFs. Both are now read.

Three planned rungs collapsed into one. A Word paragraph styled Heading 2, an h2, a tagged PDF's H2 and a layout service's section heading are one fact written four ways. ElementMetadataKeys.HeadingLevel carries it from every reader, so adding a format means teaching its reader to write that key — not writing another strategy.

The ladder is open. A host registers an IDocumentStructureStrategy rather than replacing the analyzer and losing the built-in four with it. Their orders leave gaps (100/200/300/400) so a rung slots between two without renumbering. A strategy that throws is logged and treated as having declined.

Divisions can be bounded by element, not just by page (ElementStart), which is what lets three headings on one manual page — or four stories on a newspaper page — be three and four divisions rather than one.

Advertisements left the general path. A page with no article and no running head means something in a magazine and nothing in a manual, so that split now runs only on the contents-page rung.

Measured on real documents

Running actual publications found two defects no synthetic fixture would have shown, because both come from how type is set rather than how a document is structured:

  • A wrapped headline was being read as two articles — the article split in two, the first half titled with half a sentence, and the pages handed to the half that said least. Lines now merge when they share a size, sit within ordinary leading, and overlap across the measure. All three are required: size alone swallows a standfirst, adjacency alone swallows body text, and two headlines in adjacent columns fail the overlap test.
  • Display type is routinely set twice to fake a heavier weight, so headlines read as COMMAND COMMAND & & CONQUER. Immediately-repeated words are now collapsed, in headings only.
Document Source used Result
15pp research paper (arXiv, open access) outline 7 articles + 15 sections, 3 levels deep; figures captioned
70pp magazine (CC BY-SA) inferredHeadings 39 → 27 divisions
23pp trade journal inferredHeadings 28 → 20 divisions; 69 of 103 figures captioned

No element in any of the three ended up belonging to no division — nothing would go unstored.

Breaking changes

TocSeededStructureAnalyzer is replaced by DocumentStructureAnalyzer. It was added during unreleased 2.0.0 and has never shipped, so this is its name rather than a rename. ElementMetadataKeys moves to CrestApps.Core.Abstractions, under the CrestApps.Core.Ingestion namespace. Every reader writes those keys and not every reader is about AI, so they are a contract, and a contract belongs below the things that implement it — CrestApps.Core.DataIngestion turns HTML into text and needs the heading-level key and nothing else. Its dependency closure is now itself plus one abstractions assembly, where routing it through the ingestion package had it carrying eight CrestApps assemblies, Lucene, ZString and a framework reference to ASP.NET Core to spell one constant.

Stored data needs no migration. KnowledgeObjectTypes.Section is additive.

Verification

  • 3987 tests pass; 37 new, none existing broken — including the magazine suite unchanged, which is the evidence that the multi-boundary rule was safe
  • Every project builds; the MVC host's DI container builds and serves with no errors, which a green suite never proves and which mattered here because the registration changed
  • Docs site builds clean

Known limits

Type-size inference cannot tell a masthead, a strapline or an advertisement from a headline, so a magazine's opening pages still yield divisions that are not articles. The contents-page rung handles that wherever a contents page can be read. Where overprinted runs land far enough apart to segment separately, a short fragment division can still survive.

🤖 Generated with Claude Code

MikeAlhayek and others added 9 commits September 20, 2026 10:53
The structure analyzer inferred everything from how a page looked: a
contents page, type size, where a heading sat. That is guesswork about a
layout nobody anticipated, and it is why a manual or a technical report
came out as one article -- neither carries the kind of contents page the
analyzer could read.

PdfPig has exposed the outline all along and nothing here read it.
TryGetBookmarks gives titles, nesting depth and the page each entry
opens: the document stating its own structure rather than implying it.
Virtually every user manual, technical manual, report and book carries
one, and no heuristic can beat it.

- the reader captures the outline onto the first section, because an
  IngestionDocument carries no metadata of its own and the outline
  describes the whole file rather than any one page
- the analyzer prefers it over the contents page and over type size, and
  falls through to both when a document has none, so magazines and
  newspapers are unaffected
- DocumentArticle gains Depth and ParentOrdinal, so a chapter and the
  sections beneath it are one ordered list with the nesting recorded on
  each. Everything that walks a document in reading order keeps working
  without knowing whether it is nested at all
- a division runs until the next one at its own level or shallower, so a
  chapter spans its own sections rather than stopping at the first of
  them. Deeper divisions are emitted after the ones containing them,
  which makes the innermost division own a page the two of them share and
  so keeps a chapter's text from being stored twice
- whatever precedes the first entry is front matter the outline does not
  name; it joins the first division rather than belonging to nothing,
  because an element owned by no division is one whose text is never
  stored

Design notes in docs/design/general-document-structure-engine.md.

Still to come: a section knowledge-object type, so a nested division is
stored as one. Today a depth-two division is stored as an article, which
is right about the text and wrong about the noun.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Only the PDF reader wrote anything a structure analyzer could use. Word
and HTML wrote nothing at all, so every .docx and every crawled page
ingested as one undivided article -- despite both stating their structure
more plainly than any PDF does. A paragraph styled Heading 2 and an h2
are facts, not inferences, and they were being thrown away.

One metadata key now carries that fact from every format, and one
strategy divides on it without knowing which reader produced the
document.

- ElementMetadataKeys.HeadingLevel: a one-based level an element states
- the Word reader reads it from the paragraph's outline level first and
  its Heading N style second. The outline level comes first because a
  custom style built on Heading 2 carries the level without carrying the
  name, and reading only the style would call that document unstructured
- the HTML reader now emits the blocks a page is written in rather than
  one run of its text, marking h1-h6 with their level. A page with no
  block structure still falls back to the single run it produced before
- the analyzer prefers stated levels over its contents-page and type-size
  tiers, which are both inference, and falls through to them unchanged

Divisions from stated headings are bounded by element rather than by
page, because a page carrying three headings is ordinary in a manual and
in anything converted from a word processor. DocumentArticle.ElementStart
records where a division opens, and stamping walks the elements once and
carries the open division forward. Page-bounded divisions -- an outline
entry points at a page, not a paragraph -- are untouched.

A nested division is now stored as a section rather than as another
article, hanging off the division containing it. Retrieval can be asked
for the part rather than the whole: the Methods section of a paper is a
different answer from the paper.

CrestApps.Core.DataIngestion takes a reference on the ingestion package
to reach the shared key. It is one more assembly in a graph that already
carries the AI package everywhere the HTML reader is used, and the
alternative was a second spelling of the same constant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three sources of structure that were being ignored, and the page shape
that could not be represented at all.

A tagged PDF names which of its text is a heading and at what rank. Most
government and much corporate output is tagged, because publishing under
an accessibility policy requires it. The reader now collects H1-H6 from
the page's marked content and matches them to the blocks layout analysis
produced by the text they carry -- two views of a page joined on what
they say rather than on bounds that were never meant to agree. A bare H
is taken as the outermost rank, which is what it is in the documents that
use it.

Azure Document Intelligence already reported a paragraph's role and the
mapper dropped it. Title and SectionHeading now become heading levels one
and two. This matters most for scans: a page of pixels states no type
size, so before this nothing the analyzer reads survived one, and a
scanned report divided into nothing no matter how plainly it was written.

A page carrying more than one story could not be represented. The
inference kept a single heading per page -- the topmost -- so every story
after the first on a newspaper page was absorbed into the one above it.
It now takes as many headings as a page carries, in reading order, which
is already column-wise.

The guard that made the old rule safe is kept where it earns its place:
the first heading on a page still has to sit near the top, so a pull
quote halfway down a magazine feature opens nothing. Headings after the
first are taken wherever they fall, because every headline after the
first on a newspaper page is below the fold. The existing magazine tests
pass unchanged, which is the evidence that the narrower rule was only
ever needed for the first.

Articles opened this way are bounded by element, so several can share a
page without the page's text going to whichever opened last. Rewriting an
article keeps where it opens unless its range moved, because an
element-bounded article with no opening element would hand its text to
whichever article was still open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The analyzer had grown four ways of dividing a document, tried in order,
spelled out as nested conditionals that each returned early. The order
was the important thing about the method and the hardest thing to see in
it.

The rungs are now an ordered list. Reading it top to bottom is reading
the order of authority, and adding one is a line rather than another
branch to thread a return through.

DocumentStructure.Source reports which rung answered.
DocumentStructure.IsInferred only ever said whether anything was found,
which is not the same question and not the useful one: a document divided
wrongly is far easier to argue with when the answer says whether it came
from the document's own outline, from heading levels it stated, from its
contents page, or from type size this library measured.

No behaviour changes. The two rungs that were inline -- the contents page
and the type-size inference -- became methods with the names the ladder
calls them by, and the existing tests pass unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ladder was explicit but closed: four rungs, private, reachable only
by replacing the whole analyzer and losing the other three with it. A
corpus with a convention nobody outside the business knows -- a form
series whose first line names the section, a ledger split by its own rule
-- had no way in.

IDocumentStructureStrategy is that way in. The analyzer asks every
registered strategy in Order and the first with something to say answers.
The four built-in rungs are now strategies like any other, registered the
same way, and their orders leave gaps -- 100, 200, 300, 400 -- so a host
can place one between two of them without renumbering anything.

A strategy that throws is logged and treated as having declined. One rung
failing is not the document's fault, and the rung below it may well have
an answer.

The rung implementations stayed in one internal class rather than four.
They share more than they differ -- reading a page number off a page,
deciding whether a block is a running head, measuring how far one title
is from another -- and splitting them would have meant a fifth file
holding what all four need, with every change to a shared rule landing
wherever it was least expected. What a caller sees is four small
strategies; this is the library they call into.

Advertisements left the general path on the way through. A page with no
article and no running head means something in a magazine and nothing in
a manual, so the split now runs on the table-of-contents rung rather than
on every document that reaches BuildArticles. The advertisement tests
pass unchanged, because they were always table-of-contents tests.

TocSeededStructureAnalyzer is gone; it was named after one of the four
things it had grown to do. It was added during 2.0.0 development and has
never shipped, so this is its name rather than a rename.

Also covers what was built but never asserted: that a nested division is
stored as a section, that it hangs off the division containing it, and
that its text does not also appear in its parent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two defects real publications have and synthetic fixtures never did. Both
come from how type is set rather than from how a document is structured,
which is why reasoning about the code did not find them and running
magazines through it did.

A headline that wraps is one headline. Its second line is another run of
type at the same size directly beneath the first, and the rule that lets
a newspaper page carry four stories was reading it as a fifth: the
article split in two, the first half titled with half a sentence, and the
pages handed to the half that said least.

Two lines now merge when they share a size, sit within ordinary leading
of each other, and overlap across the measure. All three conditions are
required -- same size alone would swallow the standfirst beneath a
headline, and adjacency alone would swallow the first line of body text
set large. Two headlines in adjacent columns share a size and a height
and fail the overlap test, which is what keeps them two headlines.

Display type is routinely set twice with a slight offset, to fake a
weight the font does not have. Both runs are real text, so a reader that
takes the page at its word reads every headline word twice. A word
repeated immediately after itself is now collapsed, in headings only;
body text is left alone, and English headings that legitimately repeat a
word next to themselves are vanishingly rare.

Measured on a fifteen page research paper, a seventy page magazine and a
twenty-three page trade journal. Structure came from the outline where
one existed and from type size where none did; no element in any of the
three ended up belonging to no division; figures, charts and tables were
separated, and captions were matched. The two fixes took the magazines
from thirty-nine and twenty-eight divisions to twenty-seven and twenty.

What remains is what inference cannot do: a masthead, a strapline and an
advertisement look like headlines, so the opening pages of a magazine
still yield divisions that are not articles. The contents-page rung
handles that properly wherever a contents page can be read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Earlier in this branch the HTML reader took a reference on the AI
ingestion package so it could spell one constant,
ElementMetadataKeys.HeadingLevel. That was described at the time as one
more assembly. It was not.

CrestApps.Core.DataIngestion turns HTML into text. Through that one
reference it was carrying eight CrestApps assemblies, Lucene, ZString,
Microsoft.Extensions.AI and a framework reference to ASP.NET Core. Its
dependency closure is now itself and CrestApps.Core.Abstractions, and
nothing else.

ElementMetadataKeys moves to CrestApps.Core.Abstractions, under
CrestApps.Core.Ingestion. It is a bag of const strings and it belongs
there on principle rather than convenience: every reader writes these
keys, whatever it reads and whatever it is read for, so they are a
contract, and a contract belongs below the things that implement it. A
reader should be able to state that a line is a heading without taking on
the machinery that later decides what to do about it.

The namespace is CrestApps.Core.Ingestion rather than the AI one it came
from, so the assembly and the namespace agree. The keys were added during
2.0.0 development and have never shipped, so this is where they live
rather than a move anyone has to follow.

Nothing else changed: no registration, no behaviour, and the three
readers that legitimately need the ingestion package -- PDF, OpenXml and
Document Intelligence all use MediaTypeHelper, the reader registration
helper or the outline type -- keep it. A prune pass over every file
holding a `using` for that namespace removed none, which is the evidence
that no other file was coupled to it for the keys alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@MikeAlhayek
MikeAlhayek merged commit f8cd68f into main Sep 21, 2026
9 checks passed
@MikeAlhayek
MikeAlhayek deleted the general-document-structure-engine branch September 21, 2026 14:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant