Read document structure from what a document states, not only from how it looks - #210
Merged
Merged
Conversation
The structure analyzer inferred everything from how a page looked: a contents page, type size, where a heading sat. That is guesswork about a layout nobody anticipated, and it is why a manual or a technical report came out as one article -- neither carries the kind of contents page the analyzer could read. PdfPig has exposed the outline all along and nothing here read it. TryGetBookmarks gives titles, nesting depth and the page each entry opens: the document stating its own structure rather than implying it. Virtually every user manual, technical manual, report and book carries one, and no heuristic can beat it. - the reader captures the outline onto the first section, because an IngestionDocument carries no metadata of its own and the outline describes the whole file rather than any one page - the analyzer prefers it over the contents page and over type size, and falls through to both when a document has none, so magazines and newspapers are unaffected - DocumentArticle gains Depth and ParentOrdinal, so a chapter and the sections beneath it are one ordered list with the nesting recorded on each. Everything that walks a document in reading order keeps working without knowing whether it is nested at all - a division runs until the next one at its own level or shallower, so a chapter spans its own sections rather than stopping at the first of them. Deeper divisions are emitted after the ones containing them, which makes the innermost division own a page the two of them share and so keeps a chapter's text from being stored twice - whatever precedes the first entry is front matter the outline does not name; it joins the first division rather than belonging to nothing, because an element owned by no division is one whose text is never stored Design notes in docs/design/general-document-structure-engine.md. Still to come: a section knowledge-object type, so a nested division is stored as one. Today a depth-two division is stored as an article, which is right about the text and wrong about the noun. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Only the PDF reader wrote anything a structure analyzer could use. Word and HTML wrote nothing at all, so every .docx and every crawled page ingested as one undivided article -- despite both stating their structure more plainly than any PDF does. A paragraph styled Heading 2 and an h2 are facts, not inferences, and they were being thrown away. One metadata key now carries that fact from every format, and one strategy divides on it without knowing which reader produced the document. - ElementMetadataKeys.HeadingLevel: a one-based level an element states - the Word reader reads it from the paragraph's outline level first and its Heading N style second. The outline level comes first because a custom style built on Heading 2 carries the level without carrying the name, and reading only the style would call that document unstructured - the HTML reader now emits the blocks a page is written in rather than one run of its text, marking h1-h6 with their level. A page with no block structure still falls back to the single run it produced before - the analyzer prefers stated levels over its contents-page and type-size tiers, which are both inference, and falls through to them unchanged Divisions from stated headings are bounded by element rather than by page, because a page carrying three headings is ordinary in a manual and in anything converted from a word processor. DocumentArticle.ElementStart records where a division opens, and stamping walks the elements once and carries the open division forward. Page-bounded divisions -- an outline entry points at a page, not a paragraph -- are untouched. A nested division is now stored as a section rather than as another article, hanging off the division containing it. Retrieval can be asked for the part rather than the whole: the Methods section of a paper is a different answer from the paper. CrestApps.Core.DataIngestion takes a reference on the ingestion package to reach the shared key. It is one more assembly in a graph that already carries the AI package everywhere the HTML reader is used, and the alternative was a second spelling of the same constant. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three sources of structure that were being ignored, and the page shape that could not be represented at all. A tagged PDF names which of its text is a heading and at what rank. Most government and much corporate output is tagged, because publishing under an accessibility policy requires it. The reader now collects H1-H6 from the page's marked content and matches them to the blocks layout analysis produced by the text they carry -- two views of a page joined on what they say rather than on bounds that were never meant to agree. A bare H is taken as the outermost rank, which is what it is in the documents that use it. Azure Document Intelligence already reported a paragraph's role and the mapper dropped it. Title and SectionHeading now become heading levels one and two. This matters most for scans: a page of pixels states no type size, so before this nothing the analyzer reads survived one, and a scanned report divided into nothing no matter how plainly it was written. A page carrying more than one story could not be represented. The inference kept a single heading per page -- the topmost -- so every story after the first on a newspaper page was absorbed into the one above it. It now takes as many headings as a page carries, in reading order, which is already column-wise. The guard that made the old rule safe is kept where it earns its place: the first heading on a page still has to sit near the top, so a pull quote halfway down a magazine feature opens nothing. Headings after the first are taken wherever they fall, because every headline after the first on a newspaper page is below the fold. The existing magazine tests pass unchanged, which is the evidence that the narrower rule was only ever needed for the first. Articles opened this way are bounded by element, so several can share a page without the page's text going to whichever opened last. Rewriting an article keeps where it opens unless its range moved, because an element-bounded article with no opening element would hand its text to whichever article was still open. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The analyzer had grown four ways of dividing a document, tried in order, spelled out as nested conditionals that each returned early. The order was the important thing about the method and the hardest thing to see in it. The rungs are now an ordered list. Reading it top to bottom is reading the order of authority, and adding one is a line rather than another branch to thread a return through. DocumentStructure.Source reports which rung answered. DocumentStructure.IsInferred only ever said whether anything was found, which is not the same question and not the useful one: a document divided wrongly is far easier to argue with when the answer says whether it came from the document's own outline, from heading levels it stated, from its contents page, or from type size this library measured. No behaviour changes. The two rungs that were inline -- the contents page and the type-size inference -- became methods with the names the ladder calls them by, and the existing tests pass unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ladder was explicit but closed: four rungs, private, reachable only by replacing the whole analyzer and losing the other three with it. A corpus with a convention nobody outside the business knows -- a form series whose first line names the section, a ledger split by its own rule -- had no way in. IDocumentStructureStrategy is that way in. The analyzer asks every registered strategy in Order and the first with something to say answers. The four built-in rungs are now strategies like any other, registered the same way, and their orders leave gaps -- 100, 200, 300, 400 -- so a host can place one between two of them without renumbering anything. A strategy that throws is logged and treated as having declined. One rung failing is not the document's fault, and the rung below it may well have an answer. The rung implementations stayed in one internal class rather than four. They share more than they differ -- reading a page number off a page, deciding whether a block is a running head, measuring how far one title is from another -- and splitting them would have meant a fifth file holding what all four need, with every change to a shared rule landing wherever it was least expected. What a caller sees is four small strategies; this is the library they call into. Advertisements left the general path on the way through. A page with no article and no running head means something in a magazine and nothing in a manual, so the split now runs on the table-of-contents rung rather than on every document that reaches BuildArticles. The advertisement tests pass unchanged, because they were always table-of-contents tests. TocSeededStructureAnalyzer is gone; it was named after one of the four things it had grown to do. It was added during 2.0.0 development and has never shipped, so this is its name rather than a rename. Also covers what was built but never asserted: that a nested division is stored as a section, that it hangs off the division containing it, and that its text does not also appear in its parent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two defects real publications have and synthetic fixtures never did. Both come from how type is set rather than from how a document is structured, which is why reasoning about the code did not find them and running magazines through it did. A headline that wraps is one headline. Its second line is another run of type at the same size directly beneath the first, and the rule that lets a newspaper page carry four stories was reading it as a fifth: the article split in two, the first half titled with half a sentence, and the pages handed to the half that said least. Two lines now merge when they share a size, sit within ordinary leading of each other, and overlap across the measure. All three conditions are required -- same size alone would swallow the standfirst beneath a headline, and adjacency alone would swallow the first line of body text set large. Two headlines in adjacent columns share a size and a height and fail the overlap test, which is what keeps them two headlines. Display type is routinely set twice with a slight offset, to fake a weight the font does not have. Both runs are real text, so a reader that takes the page at its word reads every headline word twice. A word repeated immediately after itself is now collapsed, in headings only; body text is left alone, and English headings that legitimately repeat a word next to themselves are vanishingly rare. Measured on a fifteen page research paper, a seventy page magazine and a twenty-three page trade journal. Structure came from the outline where one existed and from type size where none did; no element in any of the three ended up belonging to no division; figures, charts and tables were separated, and captions were matched. The two fixes took the magazines from thirty-nine and twenty-eight divisions to twenty-seven and twenty. What remains is what inference cannot do: a masthead, a strapline and an advertisement look like headlines, so the opening pages of a magazine still yield divisions that are not articles. The contents-page rung handles that properly wherever a contents page can be read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Earlier in this branch the HTML reader took a reference on the AI ingestion package so it could spell one constant, ElementMetadataKeys.HeadingLevel. That was described at the time as one more assembly. It was not. CrestApps.Core.DataIngestion turns HTML into text. Through that one reference it was carrying eight CrestApps assemblies, Lucene, ZString, Microsoft.Extensions.AI and a framework reference to ASP.NET Core. Its dependency closure is now itself and CrestApps.Core.Abstractions, and nothing else. ElementMetadataKeys moves to CrestApps.Core.Abstractions, under CrestApps.Core.Ingestion. It is a bag of const strings and it belongs there on principle rather than convenience: every reader writes these keys, whatever it reads and whatever it is read for, so they are a contract, and a contract belongs below the things that implement it. A reader should be able to state that a line is a heading without taking on the machinery that later decides what to do about it. The namespace is CrestApps.Core.Ingestion rather than the AI one it came from, so the assembly and the namespace agree. The keys were added during 2.0.0 development and have never shipped, so this is where they live rather than a move anyone has to follow. Nothing else changed: no registration, no behaviour, and the three readers that legitimately need the ingestion package -- PDF, OpenXml and Document Intelligence all use MediaTypeHelper, the reader registration helper or the outline type -- keep it. A prune pass over every file holding a `using` for that namespace removed none, which is the evidence that no other file was coupled to it for the keys alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
TocSeededStructureAnalyzerwas a magazine analyzer with a safe fallback. It read a contents page, else inferred headings from type size, else gave up and called the file one article. Three properties stopped it generalising:DocumentStructure.Articleswas flat and there was nosectiontype, so a user manual's part → chapter → procedure collapsed into one article..docxand every crawled page ingested as one undivided article — despite both stating their structure far more plainly than any PDF does.What changed
IDocumentStructureAnalyzernow runs a ladder, each rung authoritative over the next, each degrading to the one below. The contract that nothing may fail an ingest is kept: the last rung is still one division covering everything.DocumentStructureSourcesOutlineStatedHeadingsTableOfContentsInferredHeadingsWholeTwo authoritative sources shipped inside PdfPig and nothing here touched them.
TryGetBookmarksgives a hierarchical outline with page destinations — the document stating its own structure, which no heuristic can beat.Page.GetMarkedContentsgives real roles on tagged PDFs. Both are now read.Three planned rungs collapsed into one. A Word paragraph styled
Heading 2, anh2, a tagged PDF'sH2and a layout service's section heading are one fact written four ways.ElementMetadataKeys.HeadingLevelcarries it from every reader, so adding a format means teaching its reader to write that key — not writing another strategy.The ladder is open. A host registers an
IDocumentStructureStrategyrather than replacing the analyzer and losing the built-in four with it. Their orders leave gaps (100/200/300/400) so a rung slots between two without renumbering. A strategy that throws is logged and treated as having declined.Divisions can be bounded by element, not just by page (
ElementStart), which is what lets three headings on one manual page — or four stories on a newspaper page — be three and four divisions rather than one.Advertisements left the general path. A page with no article and no running head means something in a magazine and nothing in a manual, so that split now runs only on the contents-page rung.
Measured on real documents
Running actual publications found two defects no synthetic fixture would have shown, because both come from how type is set rather than how a document is structured:
COMMAND COMMAND & & CONQUER. Immediately-repeated words are now collapsed, in headings only.outlineinferredHeadingsinferredHeadingsNo element in any of the three ended up belonging to no division — nothing would go unstored.
Breaking changes
TocSeededStructureAnalyzeris replaced byDocumentStructureAnalyzer. It was added during unreleased 2.0.0 and has never shipped, so this is its name rather than a rename.ElementMetadataKeysmoves toCrestApps.Core.Abstractions, under theCrestApps.Core.Ingestionnamespace. Every reader writes those keys and not every reader is about AI, so they are a contract, and a contract belongs below the things that implement it —CrestApps.Core.DataIngestionturns HTML into text and needs the heading-level key and nothing else. Its dependency closure is now itself plus one abstractions assembly, where routing it through the ingestion package had it carrying eight CrestApps assemblies, Lucene, ZString and a framework reference to ASP.NET Core to spell one constant.Stored data needs no migration.
KnowledgeObjectTypes.Sectionis additive.Verification
Known limits
Type-size inference cannot tell a masthead, a strapline or an advertisement from a headline, so a magazine's opening pages still yield divisions that are not articles. The contents-page rung handles that wherever a contents page can be read. Where overprinted runs land far enough apart to segment separately, a short fragment division can still survive.
🤖 Generated with Claude Code