Add a system PDF agent that creates, converts, edits, previews and analyses PDFs - #224
Open
MikeAlhayek wants to merge 34 commits into
Open
MikeAlhayek wants to merge 34 commits into
MikeAlhayek wants to merge 34 commits into
Conversation
A per-conversation PDF workspace kept in the document file store, a composition model rendered by one MigraDoc composer (now also behind PdfGeneratedFileWriter), an SVG page renderer for previews drawn from the file's own geometry, and the shared tool frame the agent's tools run in.
The content rewriter removes the glyphs and images under an area from the content stream itself, keeping the surrounding text in place, and the redaction is verified by reading the result back.
…ove or cover areas A replacement removes the old glyphs from the content stream and sets the new text on the same baseline, size and colour. Numeric dates are no longer read as phone numbers, and list arguments such as categories accept a comma-separated string.
PdfModelClient resolves the text model (utility deployment, falling back to chat) and the vision model (vision slot with image input), or a chat client a host or test registers under a service key. PdfCorpus reads page texts with page numbers, chunks them for a model on page and paragraph boundaries, and splits them into overlapping passages; PdfBm25Index ranks those passages. PdfValues, PdfModelJson, PdfPageImages and PdfTextLayerWriter cover value normalization, reading a model's JSON, preparing embedded images for a vision model and writing an invisible OCR text layer.
ocr_pdf, analyze_pdf_images, summarize_pdf, ask_pdf, extract_pdf_entities, extract_pdf_data, classify_pdf and cross_reference_pdfs. Each cites pages, bounds what it sends to and returns from a model, and says plainly when no model is configured: summaries and data extraction hand the work back to the agent, while entities, classification, question answering and cross-referencing still return their pattern-based results.
Covers BM25 ranking, corpus chunking and passages, every tool's no-model answer and model path, OCR's searchable text layer read back by PdfPig, entity grouping and masking, and cross-reference conflict detection.
… marks Notes, highlights, underlines, strikeouts, squiggles, rectangles, ellipses, lines, arrows, text boxes and stamps can be placed by the text they mark or by position. Each carries its own appearance, so every viewer shows it the same way. Edits now configure the host fonts before drawing text.
The helpers read and rewrite name trees, embedded files, layers, bookmarks, signatures and streams on PDFsharp's low-level objects, and keep the XMP packet in step with the document information: PDFsharp regenerates it on every save without escaping values and drops PDF/A and PDF/UA ids, so the packet it writes last is replaced in place with a merged, escaped one.
Lists, adds, replaces and clears a PDF's outline, and builds one from the headings: lines drawn clearly above the body size, ranked into levels by size, with running heads, page numbers and a one-off title left out.
Lists, embeds, extracts and removes embedded files and page file attachments. The embedded-files tree is read and rebuilt here because PDFsharp's AddEmbeddedFile replaces an opened document's existing tree; files whose type can run code are downloaded as octet-stream.
Lists optional content groups with their default visibility, lock and the pages whose resources use them, and shows or hides layers by moving them between the default configuration's ON and OFF arrays.
Places web links (http, https and mailto only) and links to pages over text found with PdfTextFinder or over an area given in top-left points; PDFsharp link rectangles are plain user space, as the tests read back.
protect_pdf encrypts with AES-256 (or AES-128, RC4-128), an open password and permissions; without an owner password for restricted permissions a random one is generated and never shown. remove_pdf_security only takes a password the user supplied. PDFsharp decrypts documents as it opens them, so encryption is read from the file, and the editing tools now say when a protected source's working copy is saved without its protection.
sign_pdf signs with the host's IPdfSigningCertificateProvider through a CMS signer that signs the message digest, signing time and ESS signing-certificate-v2 (PDFsharp's default signer signs none of them), refuses a PDF that is already signed, and checks its own result. verify_pdf_signature hashes the byte ranges, checks the PKCS#7 signature and message digest, the certificate chain and whether bytes follow the signed revision.
Removes page thumbnails and, when asked or at level maximum, metadata; merges reachable streams that are byte-for-byte identical and rewrites references to one copy; compresses unfiltered streams and, at maximum, recompresses Flate streams at the best level. It reports the size before and after and keeps the original when the result is not smaller.
Twelve tools that read PDFs, registered in the reading group: extract_pdf_text, extract_pdf_tables, extract_pdf_images, extract_pdf_structure, extract_pdf_links, search_pdf, get_pdf_page_content, analyze_pdf_layout, detect_pdf_language, generate_pdf_outline, compare_pdfs and convert_from_pdf. Shared helpers live under Analysis/Reading: a layout analyzer that assigns block roles (headings by type-size tiers, running heads and feet, captions, list items, table and figure text) and columns, a structure builder and writers (outline, Markdown, HTML, JSON), table extraction with cross-page merging through the ingestion reader, a link reader that resolves internal targets and named destinations without following any link, an offline language detector, and a linear-space Myers diff for comparisons.
Covers every reading tool against PDFs built in code (a composed report with headings, a list, a table and a captioned picture; a long table that runs over two pages; a two-column PDFsharp document with running heads and page numbers; links, a sticky note and bookmarks; a password-protected copy), plus unit tests for the language detector, the text diff, URI validation and the registrations.
…opped Every edit now takes the file's XMP packet when it opens the document and writes an escaped packet that keeps the PDF/A and PDF/UA identification when it saves, instead of PDFsharp's stale copy. An answer notes when a protected source was saved without its protection, and a read-only open no longer asks for the owner password.
A safe content-stream parser (PDFsharp's reader exhausts memory on an unterminated string), a scanner for the resources and fonts content names, a font inventory, structure tree, form field, annotation and active-content readers, XMP reading and preservation, and the check modules for integrity, rendering, overflow, fonts, links, layout, composed-definition comparison, accessibility and PDF/A / PDF/UA.
validate_pdf, check_pdf_quality, check_pdf_accessibility, tag_pdf_accessibility and validate_pdf_compliance, registered in the quality tool group.
Fixture PDFs are built in code: composed documents, PDFsharp pages with hand-written content, and raw files that carry exact defects.
…rotected files PDFsharp's PDF/A switch also turns on a structure mode MigraDoc cannot draw in, so composing with pdf_a failed. The composer now adds the sRGB output intent and writes the PDF/A identification into escaped XMP itself. The quality helpers' object reader and tag element are renamed so they no longer collide with the properties and reading helpers, and preview_pdf draws a protected upload from a decrypted copy.
…ntent By default it removes what the pages do not show: document information and XMP, scripts and automatic actions, attached files, thumbnails, invisible text and hidden layers with their content. Comments, form data and external links are removed when named, and the answer says which of them remain.
…nto a composed PDF Word documents keep their headings, formatted text, links, lists, tables and pictures; decks become a landscape page per slide with optional speaker notes; sheets read through the tabular workspace keep their formats, with a direct read as the fallback; Markdown, HTML, text and images convert too, and several files become one PDF. A tool that adds an asset can now use it in the same call.
The AI Documents guide gains a section on the PDF Agent — its workspace, tools, previews, generated PDFs, options and signing — the agents guide lists it beside the Tabular Data Agent, and the 2.0.0 notes describe it with the XMP and PDF/A fixes and the new public API.
The Linux font resolver served one regular font for every family and style and had PDFsharp fake bold and italic, so a PDF composed on Linux had no real bold, italic, serif or monospaced text. It now serves the installed family, or Liberation Sans, Serif or Mono in its place, with the face for each style, and simulates a style only when the font has no face for it. The PDF test fixtures relied on PDFsharp's default page size, which is Letter on Windows and A4 on the Linux runner, and on Arial being installed; they now set the page size and match fonts by style.
A second caller used to return as soon as the first had set the flag, before the Linux font resolver was installed, and drew text with no resolver.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #223.
This adds the system PDF Agent (
pdf-agent), built the same way as the Tabular Data Agent. It is anIAIProfileProvideragent with these properties:AlwaysAvailableandIsSystem, so the primary model can always delegate to it.pdf-agent.mdtemplate.It runs 53 hidden tools over a per-conversation PDF workspace.
AddPdf()registers it, andPdfAgentOptions.Enabled = falseturns it off.Workspace
contract.pdfsaves a working copy namedcontract, and later edits change that copy.IDocumentFileStore. It is removed when the conversation's documents are cleaned up, through the newIConversationWorkspaceCleanupHandler, and when chat history is cleared.Tools
create_pdf,add_pdf_content,format_pdf,preview_pdf,export_pdf,get_pdf_infoconvert_to_pdf(Word, PowerPoint, spreadsheets/CSV, Markdown, HTML, text, images; several files into one PDF),convert_from_pdfedit_pdf_pages(merge, split, extract, reorder, reverse, rotate, delete, duplicate, insert blank, crop, resize, watermark, stamp, page numbers, header/footer)edit_pdf_content,manage_pdf_annotations,add_pdf_bookmarks,add_pdf_links,edit_pdf_metadata,manage_pdf_attachments,manage_pdf_layers,flatten_pdfget_pdf_form_fields,fill_pdf_form,edit_pdf_form,validate_pdf_formredact_pdf,find_pdf_sensitive_data,sanitize_pdf,protect_pdf,remove_pdf_security,sign_pdf,verify_pdf_signatureextract_pdf_text,extract_pdf_tables,extract_pdf_images,extract_pdf_structure,extract_pdf_links,search_pdf,get_pdf_page_content,analyze_pdf_layout,detect_pdf_language,generate_pdf_outline,compare_pdfsocr_pdf,analyze_pdf_images,summarize_pdf,ask_pdf,extract_pdf_entities,extract_pdf_data,classify_pdf,cross_reference_pdfsvalidate_pdf,check_pdf_quality,check_pdf_accessibility,tag_pdf_accessibility,validate_pdf_compliance,optimize_pdfThe agent can also call
list_tabular_dataandquery_tabular_data, so table and chart blocks can read from uploaded spreadsheets.Mapping of the tools in @MikeAlhayek's comment
Some of the requested tools were merged into one tool with an action or operation argument. That keeps the agent's tool list shorter for the model:
add_pdf_toc,add_pdf_header_footer,add_pdf_page_numbers,add_pdf_watermark,set_pdf_page_layoutformat_pdffor composed PDFs;edit_pdf_pages(header_footer,page_numbers,watermark,resize,rotate) for any PDFadd_pdf_table,add_pdf_chart,add_pdf_image,add_pdf_pageadd_pdf_contentblocks (table,chart,image,page_break);edit_pdf_content(add_image) andedit_pdf_pages(insert_blank) for existing PDFsextract_pdf_annotations,add_pdf_annotations,edit_pdf_annotationsmanage_pdf_annotations(list/add/update/remove)create_pdf_form,add_pdf_signature_fieldedit_pdf_form(text, multiline, password, checkbox, radio, dropdown, list box and signature fields)flatten_pdf_formflatten_pdf(fields, annotations or both)compress_pdfoptimize_pdfrender_pdf_pagespreview_pdf(in chat), plusconvert_from_pdfwithsvg(as files)detect_pdf_tablesextract_pdf_tablesget_pdf_reading_orderanalyze_pdf_layoutcheck_pdf_rendering,check_pdf_fonts,check_pdf_links,check_pdf_content_overflow,validate_pdf_layoutcheck_pdf_quality(one check group per concern)Every other requested name exists as a tool of the same name.
How the main pieces work
preview_pdfhas no rasterizer to call, so it draws each page as SVG from the page's own geometry: text runs, paths, images, annotations and form values. It works the same for uploads and composed documents. Previews go throughIGeneratedDocumentServiceand come back as[fig:N]markers, like the tabular preview. Output is bounded, and the tool falls back to the page text when the host cannot serve pictures. As with the tabular preview, the preview format is resolvable but not requestable.generate_fileandexport_tabular_data. A tabular export as PDF keeps what the tabular preview shows: number formats, header colours, the total row, highlighted cells and charts. Each sheet gets its own page, and wide sheets are landscape.redact_pdfrewrites the content stream to remove the glyphs and images under each area, keeping the neighbouring text in place. It then reads the file back to confirm the text is gone.sign_pdfonly signs with a certificate the host configures, throughAddPdfSigning()or anIPdfSigningCertificateProvider. It never uses one supplied in the conversation. It uses a CMS signer that signs the message digest, the signing time and the signer-certificate hash.Fixes that apply to every PDF the framework writes
&broke the packet. It also drops the PDF/A and PDF/UA identification and leaves the old packet in the file. Every save now writes an escaped packet that keeps the identification.pdf_athrew. PDFsharp's PDF/A switch also turns on a PDF/UA structure mode that MigraDoc cannot draw in. The composer now adds the sRGB output intent and the identification itself. The result passesvalidate_pdf_compliancefor PDF/A-2b.2026-03-15was classed as a phone number. It is now a date.New public API
IConversationWorkspaceCleanupHandlerGeneratedFileContent.EncodedContentandGeneratedDocumentRequest.ContentTypePdfAgentOptions,PdfCompositionOptions,PdfPreviewOptionsPdfSigningOptions,IPdfSigningCertificateProvider,AddPdfSigning()DefaultConversationDocumentCleanupServicegains a constructor that takes the handlers; the old constructor still works. The PDF project now referencesDocumentFormat.OpenXml, which the conversion tools need.Known limits
convert_to_pdfsets the source's content in the PDF theme; it does not reproduce the original layout. Text boxes, charts, equations, footnotes and Word headers/footers are reported as not converted./Rotateentry are in the page's unrotated frame.Acceptance criteria
PdfTabularExportTests).tests/.../Core/Documents/Pdf/**.core/ai-documents.md(new PDF files and the PDF Agent section) andcore/agents.md. Changelog:changelog/2.0.0.md.Testing
🤖 Generated with Claude Code