Document engineering | OCR pipelines | Zotero automation | local AI tooling
I build local-first tools for converting scanned documents, PDFs, and academic collections into searchable, structured, reproducible knowledge assets.
I work on practical document-engineering systems: OCR pipelines, searchable PDFs, academic-library automation, and local-first AI tooling. Most of my projects are built around one idea: keep the workflow reproducible, inspectable, and useful on a local machine before adding external services.
- I like tools that turn messy document collections into structured knowledge.
- I care about OCR geometry, hOCR, PDF assembly, metadata, and traceability.
- I build Windows-friendly GUI/CLI utilities for workflows that usually live as fragile scripts.
- I am exploring local LLM workflows for translation, review, orchestration, and RAG-ready document preparation.
| Domain | Engineering output |
|---|---|
| OCR and searchable PDFs | Local OCR pipelines, hOCR bridges, OCRmyPDF plugins, searchable PDF assembly. |
| PDF processing | Page cleanup, geometry correction, crop benchmarking, image-to-PDF workflows. |
| Zotero workflows | Local Zotero storage processing, PDF-to-Markdown conversion, WebDAV mirroring, traceability artifacts. |
| Local AI | LM Studio workflows, local model routing, document translation experiments, self-hosted tooling. |
| Automation | Windows-first GUI/CLI tools, repeatable setup scripts, workflow orchestration. |
| Information extraction | RAG-ready document preparation and structured document conversion. |
| Project | Problem domain | Role in the ecosystem | Status |
|---|---|---|---|
| Surya_Chandra_PDF_OCR | OCR, searchable PDFs, document AI | Core local OCR pipeline for turning scanned PDFs into searchable PDF assets. | Core project |
| ZoteroPDF_2_MD | Zotero, Marker, academic workflows | Processes local Zotero PDF collections into Markdown and HTML workflows with traceability. | Active workflow tool |
| img_2_pdf | PDF preprocessing, document geometry | Pre-OCR layer for page capture, cleanup, crop, perspective correction, and PDF export. | Foundation layer |
| Claude Code Local LLM Switcher | Local AI, LM Studio, developer tooling | Windows GUI/CLI for mapping Claude Code aliases to local LM Studio models and isolated VS Code sessions. | Local AI utility |
| Layer | Repository | Purpose |
|---|---|---|
| OCR pipeline | Surya_Chandra_PDF_OCR |
Local scanned-PDF to searchable-PDF workflow. |
| Geometry bridge | surya_hOCR_bridge |
Converts Surya sidecars and OCR artifacts into hOCR/searchable-PDF workflows. |
| OCRmyPDF plugin | Surya_OCRmyPDF_Plugin |
Surya OCR integration for OCRmyPDF. |
| OCRmyPDF plugin | Chandra_OCRmyPDF_Plugin |
Chandra layout OCR to hOCR/OCRmyPDF bridge. |
| Plugin hub | OCRmyPDF_Plugins_HUB |
Working snapshots and integration references for OCRmyPDF plugin experiments. |
img_2_pdf is the pre-OCR document pipeline layer:
- page import and capture;
- deskew, crop, perspective correction, and enhancement;
- image export and merged PDF generation;
- crop benchmark tooling;
- clean separation from OCR benchmarking and plugin repositories.
Related repositories:
Surya_Chandra_PDF_OCR- searchable PDF assembly;surya_hOCR_bridge- OCR geometry and hOCR bridge;OCRmyPDF_Plugins_HUB- OCRmyPDF plugin integration workspace;mrf.museumart.ru-pdf_book_saver- book/PDF saving automation.
ZoteroPDF_2_MD connects local academic collections with document conversion workflows:
- local Zotero PDF collection processing;
- Markdown and HTML conversion paths;
- deterministic staging and traceability artifacts;
- WebDAV-oriented file mirroring;
- local model experiments for document translation and review.
Zotero_Integration_Metadata_Mapping and Zotero_SciHub_module are adjacent experiments around metadata and academic-source automation.
| Repository | Focus | Notes |
|---|---|---|
Claude_Code_Local_LLM_Switcher_LM_Studio |
Claude Code and LM Studio | Alias binding, local endpoint checks, context control, isolated VS Code launcher. |
ZoteroPDF_2_MD |
Marker and local model experiments | Academic PDF conversion, HTML workflows, local translation experiments. |
LLM_Task_Orchestration |
Task orchestration | Experimental workspace for future local AI workflow coordination. |
LLM_translator_test |
Translation experiments | Early local translation and workflow testing. |
Secondary work includes GUI wrappers, local workflow launchers, book/PDF savers, OCR plugin experiments, document conversion utilities, and workflow orchestration prototypes.
| Repository | Area |
|---|---|
ISO_Sensitometer_GUI |
Imaging and measurement GUI tooling. |
Obsidian_Doctor |
Local knowledge-base maintenance. |
marker_GUI |
GUI experimentation around document conversion. |
Languages and runtime: Python, PowerShell, Bash, Docker
Document tooling: OCRmyPDF, hOCR, pypdf, pypdfium2, ReportLab, Marker
AI/OCR: Surya, Chandra, LM Studio, local model workflows
Data and workflow: SQLite, Zotero storage, WebDAV, CSV/JSON traceability artifacts
Interfaces: CLI, Tkinter GUI, local HTTP services, Windows automation
If these tools save you time, support links can be added here once the public accounts are connected.
- GitHub: @NixWrk
- International identity: add LinkedIn and ORCID once public profiles are ready.
- Best project contact path: open an issue in the relevant repository.
- Collaboration focus: OCR/PDF automation, Zotero workflows, local AI tooling, information extraction, and RAG-ready document pipelines.

