Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
2 changes: 1 addition & 1 deletion Document-Processing/Common/font-manager.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
title: Syncfusion font handling in Office-to-PDF and image conversions
title: Font Manager for Office-to-PDF or image conversion | Syncfusion
description: Learn how Syncfusion Document Processing handles font management during Office to PDF/Image conversions and PDF processing workflows.
platform: document-processing
documentation: UG
Expand Down
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
---
title: Assemblies required for Data Extraction | Syncfusion
description: This section details the Syncfusion assemblies required to configure and run Data Extraction seamlessly in .NET projects.
title: Assemblies Required in .NET Smart Data Extractor | Syncfusion
description: This section describes the required Syncfusion assemblies needed to integrate and use the Smart Data Extractor effectively in your applications
platform: document-processing
control: DataExtraction
documentation: UG
keywords: Assemblies
---
# Assemblies required for Data Extraction
# Assemblies Required in Data Extraction

## Smart Data Extractor

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Extract data in ASP.NET Core | Syncfusion
title: Getting Started with ASP.NET Core Smart Data Extractor | Syncfusion
canonical_url: "https://www.syncfusioninternal.com/document-sdk/net-pdf-data-extraction"
description: Learn how to extract data from PDF in ASP.NET Core with step‑by‑step guidance using Syncfusion .NET Core Data extraction library.
description: Learn how to get started with the Syncfusion ASP.NET Core Smart Data Extractor. Explore setup, features, examples, and customization options.
platform: document-processing
control: SmartDataExtractor
documentation: UG
---

# Extract Data in ASP.NET Core
# Getting Started with ASP.NET Core Smart Data Extractor

The Syncfusion<sup>&reg;</sup> Smart Data Extractor is a .NET library used to extract structured data and document elements from PDF and image files in ASP.NET Core applications.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
---
title: Extract Data in ASP.NET MVC Application | Syncfusion
description: Learn how to extract data in an ASP.NET MVC application with step‑by‑step guidance using the Syncfusion Data Extraction library.
title: Getting Started with ASP.NET MVC Smart Data Extractor | Syncfusion
description: Learn how to get started with the Syncfusion ASP.NET MVC Smart Data Extractor. Explore setup, features, examples, and customization options.
platform: document-processing
control: SmartDataExtractor
documentation: UG
keywords: Assemblies

---

# Extract Data in ASP.NET MVC
# Getting Started with ASP.NET MVC Smart Data Extractor

The Syncfusion<sup>&reg;</sup> Smart Data Extractor is a .NET library used to extract structured data and document elements from PDFs and images in ASP.NET MVC applications.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
---
title: Extract Data in Blazor Application | Syncfusion
description: Learn to extract tables, forms, text, and images from PDF documents and scanned images in Blazor using the Syncfusion® Smart Data Extractor .NET library.
title: Getting Started with Blazor Smart Data Extractor | Syncfusion
description: Learn how to get started with the Syncfusion Blazor Smart Data Extractor. Explore setup, features, examples, and customization options.
platform: document-processing
control: SmartDataExtractor
documentation: UG
keywords: Assemblies

---

# Extract Data from PDF in Blazor
# Getting Started with Blazor Smart Data Extractor

The Syncfusion<sup>&reg;</sup> Smart Data Extractor is a .NET library used to extract structured data and document elements from PDFs and images in Blazor applications.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
---
title: Extract Data in Console Application | Syncfusion
description: Learn how to extract data in a Console Application by using the .NET Smart Data Extractor Library efficiently.
title: Getting Started with Console Smart Data Extractor | Syncfusion
description: Learn how to get started with the Syncfusion Console Smart Data Extractor. Explore setup, features, examples, and customization options.
platform: document-processing
control: SmartDataExtractor
documentation: UG
---

# Extract Data from PDF in Console Application
# Getting Started with Console Smart Data Extractor

The Syncfusion<sup>&reg;</sup> Smart Data Extractor is a .NET library used to extract structured data and document elements from PDFs and images in Console applications.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
---
title: Extract Data in .NET MAUI | Syncfusion
description: Extract tables, forms, text, and images from PDF documents and scanned files in .NET MAUI using the Syncfusion® Smart Data Extractor.
title: Getting Started with .NET MAUI Smart Data Extractor| Syncfusion
description: Learn how to get started with the Syncfusion .NET MAUI Smart Data Extractor. Explore setup, features, examples, and customization options.
platform: document-processing
control: SmartDataExtractor
documentation: UG
keywords: Assemblies

---

# Extract Data from PDF in .NET MAUI
# Getting Started with .NET MAUI Smart Data Extractor

The Syncfusion<sup>&reg;</sup> Smart Data Extractor is a .NET library used to extract structured data and document elements from PDFs and images in .NET MAUI applications.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
---
title: Extract Data in WPF Application | Syncfusion
description: Learn how to extract data in a WPF application with step‑by‑step guidance using the .NET Smart Data Extractor Library.
title: Getting Started with WPF Smart Data Extractor | Syncfusion
description: Learn how to get started with the Syncfusion WPF Smart Data Extractor. Explore setup, features, examples, and customization options.
platform: document-processing
control: SmartDataExtractor
documentation: UG
keywords: Assemblies

---

# Extract Data from PDF in WPF
# Getting Started with WPF Smart Data Extractor

The Syncfusion<sup>&reg;</sup> Smart Data Extractor is a .NET library used to extract structured data and document elements from PDFs and images in WPF applications.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Extract Data from PDF in Windows Forms | Syncfusion
description: Extract tables, text, and form fields from PDF documents in Windows Forms using the .NET Smart Data Extractor Library.
title: Getting Started with Windows Forms Smart Data Extractor | Syncfusion
description: Learn how to get started with the Syncfusion Windows Forms Smart Data Extractor. Explore setup, features, examples, and customization options.
platform: document-processing
control: SmartDataExtractor
documentation: UG

---

# Extract Data in Windows Forms
# Getting Started with Windows Forms Smart Data Extractor

The Syncfusion<sup>&reg;</sup> Smart Data Extractor is a .NET library used to extract structured data and document elements from PDFs and images in Windows Forms applications.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: NuGet Packages for Data Extraction | Syncfusion&reg;
description: Learn the NuGet packages required to use Syncfusion&reg; Data Extraction in various platforms and frameworks.
title: NuGet Packages required for .NET Smart Data Extractor | Syncfusion
description: Discover the NuGet packages required to integrate Smart Data Extractor across .NET platforms and frameworks
platform: document-processing
control: DataExtraction
documentation: UG
keywords: Assemblies
---

# NuGet Packages Required for Data Extraction
# NuGet Packages required for .NET Smart Data Extractor

## Smart Data Extractor

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Smart Data Extractor Library | Syncfusion
description: Smart Data Extractor converts PDF documents and images to structured formats like JSON, Markdown (MD), and PDF output.
title: About Document Conversions of .NET Smart Data Extractor | Syncfusion
description: Learn about overview of the document conversions supported by Syncfusion .NET Smart Data Extractor and more details.
platform: document-processing
control: SmartDataExtractor
documentation: UG
keywords: SmartDataExtractor, PDF to JSON, PDF to Markdown
---

# Welcome to Syncfusion<sup>&reg;</sup> Smart Data Extractor Library
# About Document Conversions of .NET Smart Data Extractor

Syncfusion<sup>&reg;</sup> Smart Data Extractor Library extracts structured information from PDF documents and scanned images. It supports conversions such as **PDF to JSON**, **PDF to Markdown (MD)**, and generating **PDF output** by analyzing visual layout patterns like text blocks, tables, headers, and form fields. This helps developers easily integrate the extractor to achieve required data conversions while focusing on the core logic of their applications.

Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Extract PDF to JSON in C# | Smart Data Extractor | Syncfusion
description: Learn how to extract structured data from PDF documents as JSON in C# using the Syncfusion® Smart Data Extractor library for .NET applications.
title: Convert PDF to JSON in .NET Smart Data Extractor | Syncfusion
description: Extract structured data from PDF documents as JSON using Smart Data Extractor. Convert PDF content into machine-readable JSON format seamlessly in .NET applications.
platform: document-processing
control: SmartDataExtractor
documentation: UG
keywords: Assemblies
---

# PDF to JSON Extraction
# Convert PDF to JSON in .NET Smart Data Extractor

JavaScript Object Notation (JSON) is a lightweight data‑interchange format that is easy for humans to read and write, and simple for machines to parse and generate. The Syncfusion<sup>&reg;</sup> Smart Data Extractor library extracts structured information from PDF documents and scanned images, and outputs the content as JSON. It analyzes text blocks, tables, headers, and form fields to preserve structure, enabling developers to integrate PDF to JSON extraction into their applications.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Extract PDF to Markdown in C# | Smart Data Extractor | Syncfusion
description: Extract PDF documents as Markdown (MD) in C# using Syncfusion<sup>&reg;</sup> Smart Data Extractor library without Microsoft Office or Adobe dependencies
title: Convert PDF to Markdown in .NET Smart Data Extractor | Syncfusion
description: Extract PDF documents as Markdown using Smart Data Extractor. Convert PDF content into clean, structured Markdown content in .NET.
platform: document-processing
control: SmartDataExtractor
documentation: UG
keywords: Assemblies
---

# PDF to Markdown Extraction
# Convert PDF to Markdown in .NET Smart Data Extractor

Markdown is a lightweight markup language that adds formatting elements to plain text documents. The Syncfusion<sup>&reg;</sup> Smart Data Extractor library extracts structured information from PDF documents and scanned images, and outputs the content as Markdown (MD). It analyzes text blocks, tables, headers, and form fields to preserve layout and formatting.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
---
title: Data Extraction and Conversion in .NET | Syncfusion
title: About Syncfusion .NET Smart Data Extractor | Syncfusion
canonical_url: "https://www.syncfusioninternal.com/document-sdk/net-pdf-data-extraction"
description: Syncfusion Data Extraction is a .NET library that extracts tables, forms, text, and images from PDF or image files, and outputs JSON or Markdown.
description: Learn about introduction of Syncfusion .NET Smart Data Extractor for extracting data from PDFs or scanned images and more details.
platform: document-processing
control: DataExtraction
documentation: UG
keywords: Assemblies
---

# Welcome to .NET Smart Data Extractor Library
# About Syncfusion .NET Smart Data Extractor

{% doccards %}

Expand Down
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
---
layout: post
title: Installing Syncfusion Data Extraction - Syncfusion
description: Learn how to install the .NET Smart Data Extractor Library for extracting structured data from PDFs and images in .NET applications.
title: How to install .NET Smart Data Extractor Add-on | Syncfusion
description: Install the .NET Smart Data Extractor Add-on with this step-by-step guide
platform: document-processing
control: Installation and Deployment
documentation: ug

---

# Download Syncfusion<sup>&reg;</sup> Data Extraction Add-On
# How to install .NET Smart Data Extractor Add-on

The Syncfusion<sup>&reg;</sup> Data Extraction Add-On can be downloaded from the [Syncfusion download page](https://www.syncfusion.com/downloads).

Expand Down
6 changes: 3 additions & 3 deletions Document-Processing/Data-Extraction/NET/ocr-overview.md
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
---
title: Intro to OCR Processor | Syncfusion
title: About Syncfusion OCR Processing Library | Syncfusion
canonical_url: "https://www.syncfusion.com/document-sdk/net-pdf-library/ocr-process"
description: This page introduces the Syncfusion OCR Processor, highlighting its purpose, main features, and how to begin optical character recognition in .NET apps.
description: Learn about introduction of Syncfusion OCR Processor for recognizing text from scanned images and more details.
platform: document-processing
control: OCRProcessor
documentation: UG
keywords: OCR, Optical Character Recognition, Text Recognition
---

# Welcome to Syncfusion OCR Processor Library
# About Syncfusion OCR Processing Library

{% doccards %}

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Perform OCR on PDF and image files in AWS Textract | Syncfusion
description: Learn how to perform OCR on scanned PDF documents and images in AWS Textract using Syncfusion .NET OCR library.
title: Getting Started with AWS Textract OCR Processor | Syncfusion
description: Learn how to get started with the Syncfusion AWS Textract OCR Processor. Explore setup, features, examples, and customization options.
platform: document-processing
control: PDF
documentation: UG
keywords: Assemblies
---

# Perform OCR with AWS Textract
# Getting Started with AWS Textract OCR Processor

The [.NET OCR library](https://www.syncfusion.com/document-sdk/net-pdf-library/ocr-process) supports external OCR engines such as AWS Textract to process OCR on images and PDF documents.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: .NET 8 & Tesseract OCR on Amazon Linux 2023 EC2 | Syncfusion
description: Install & configure .NET 8, Tesseract OCR on Amazon Linux 2023 EC2 to perform OCR on PDFs & images using Syncfusion .NET OCR library.
title: Getting Started with Amazon Linux EC2 OCR Processor | Syncfusion
description: Learn how to get started with the Syncfusion Amazon Linux 2023 EC2 OCR Processor. Explore setup, features, examples, and customization options.
platform: document-processing
control: PDF
documentation: UG
keywords: Assemblies
---

# Perform OCR with Tesseract on Amazon Linux EC2 using .NET application
# Getting Started with Amazon Linux EC2 OCR Processor

The [.NET OCR library](https://www.syncfusion.com/document-sdk/net-pdf-library/ocr-process) is used to extract text from scanned PDFs and images in Linux applications with the help of Google's [Tesseract](https://github.com/tesseract-ocr/tesseract) Optical Character Recognition engine.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
---
title: Assemblies Required for OCR | Syncfusion
title: Assemblies Required in .NET OCR Processor | Syncfusion
description: This section describes the required Syncfusion assemblies needed to integrate and use the OCR Processor effectively in your applications
platform: document-processing
control: PDF
documentation: UG
keywords: Assemblies
---
# Assemblies Required to work with OCR processor
# Assemblies Required in .NET OCR Processor

Get the following required assemblies by downloading the OCR library installer. Download and install the OCR library for Windows, Linux, and Mac respectively. Please refer to the advanced installation steps for more details.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Deploy and manage with Azure Kubernetes Service | Syncfusion
description: Learn how to deploy, scale, and manage containerized applications in Azure using Azure Kubernetes Service
title: Getting Started with Azure Kubernetes Service OCR Processor| Syncfusion
description: Learn how to get started with the Syncfusion Azure Kubernetes Service OCR Processor. Explore setup, features, examples, and customization options.
platform: document-processing
control: PDF
documentation: UG
keywords: Assemblies
---

# Perform OCR with Azure Kubernetes Service
# Getting Started with Azure Kubernetes Service OCR Processor

The [.NET OCR library](https://www.syncfusion.com/document-sdk/net-pdf-library/ocr-process) can be integrated with external OCR engines like Azure Computer Vision and deployed on Azure Kubernetes Service (AKS) to efficiently process OCR tasks on images and PDF documents at scale.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Perform OCR on PDF and image files in Azure Vision | Syncfusion
description: Learn how to perform OCR on scanned PDF documents and images in Azure Vision using Syncfusion .NET OCR library.
title: Getting Started with Azure Vision OCR Processor| Syncfusion
description: Learn how to get started with the Syncfusion Azure Vision OCR Processor. Explore setup, features, examples, and customization options.
platform: document-processing
control: PDF
documentation: UG
keywords: Assemblies
---

# Perform OCR with Azure Vision
# Getting Started with Azure Vision OCR Processor

The [.NET OCR library](https://www.syncfusion.com/document-sdk/net-pdf-library/ocr-process) supports external OCR engines such as Azure Computer Vision to process OCR on images and PDF documents.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
---
title: Perform OCR on PDF and image files in Console | Syncfusion
description: Learn how to perform OCR on scanned PDF documents and images with different tesseract versions in a Console App by using the Syncfusion PDF library efficiently
title: Getting Started with Console OCR Processor| Syncfusion
description: Learn how to get started with the Syncfusion Console OCR Processor. Explore setup, features, examples, and customization options.
platform: document-processing
control: PDF
documentation: UG
---

# Perform OCR in Console Application
# Getting Started with Console OCR Processor

The [.NET OCR library](https://www.syncfusion.com/document-sdk/net-pdf-library/ocr-process) is used to extract text from scanned PDFs and images in console applications with the help of Google's [Tesseract](https://github.com/tesseract-ocr/tesseract) Optical Character Recognition engine.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
---
title: Perform OCR on PDF and image files in Docker | Syncfusion
description: Learn how to perform OCR on scanned PDF documents and images in Docker with different tesseract versions using Syncfusion .NET OCR library.
title: Getting Started with Docker OCR Processor| Syncfusion
description: Learn how to get started with the Syncfusion Docker OCR Processor. Explore setup, features, examples, and customization options.
platform: document-processing
control: PDF
documentation: UG
keywords: Assemblies
---
# Perform OCR in Docker
# Getting Started with Docker OCR Processor

The [.NET OCR library](https://www.syncfusion.com/document-sdk/net-pdf-library/ocr-process) is used to extract text from scanned PDFs and images in Docker applications with the help of Google's [Tesseract](https://github.com/tesseract-ocr/tesseract) Optical Character Recognition engine.

Expand Down
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
---
title: Perform OCR on PDF and image files | Syncfusion
description: Learn how to perform OCR on scanned PDF documents and images with different tesseract version using Syncfusion .NET OCR library.
title: Perform OCR on PDF and image files in .NET | Syncfusion
description: Learn how to perform OCR on scanned PDF documents and images with different tesseract version using Syncfusion .NET OCR Processor.
platform: document-processing
control: PDF
documentation: UG
keywords: Assemblies
---

# OCR Processor Features
# Perform OCR on PDF and image files in .NET

## Performing OCR for an entire document

Expand Down
Loading