Manages the OCR documents of this IOcrEngine.
public interface IOcrDocumentManager Public Interface IOcrDocumentManager public interface IOcrDocumentManager @interface LTOcrDocumentManager : NSObject public class OcrDocumentManager function Leadtools.Forms.Ocr.IOcrDocumentManager() public interface class IOcrDocumentManager You can access the instance of the IOcrDocumentManager used by an IOcrEngine through the IOcrEngine.DocumentManager property.
The IOcrDocumentManager interface allows you to create IOcrDocument objects that encapsulate an OCR'ed document. Each OCR document contains a collection of IOcrPage that you can use to add and remove pages from the document. After you add the pages to the document and optionally manage the zones on the pages, you can call the IOcrPage.Recognize method on each page to obtain the recognition data and store them internally in the pages. Once you are done, you can use the save methods of the IOcrDocument object to save the document into its final format.
LEADTOOLS supports saving to various standard document formats such as PDF, Microsoft Word, HTML and several others through the LEADTOOLS Document Writers engine. For more information, refer to IOcrDocument and DocumentFormat.
IOcrDocumentManager supports creating documents in two ways:
In this mode, the OCR pages are required to be in memory before saving. This is not recommended when the document have a large amount of pages and either using a file-based document or using the LEADTOOLS Temporary file format (DocumentFormat.Ltd is required.
In memory-based IOcrDocument, the IOcrPageCollection holds the pages. The user can recognize any or all of the pages at any time and pages can be added or removed at will.
Use or IOcrDocumentManager.CreateDocument or IOcrDocumentManager.CreateDocument(string, OcrCreateDocumentOptions) with the OcrCreateDocumentOptions.InMemory flag to create such documents.
IOcrDocument.IsInMemory will be true for memory-based documents.
In this mode, the OCR pages are not required to be in memory before saving. This mode is recommended when the document have a large amount of pages and.
In file-based IOcrDocument, the IOcrPageCollection is a store-only view of the pages. when page is added, a snap shot of the current recognition data is saved into the document. This data cannot be modified any more and the page is no longer needed. The user must recognize the pages before they are added to the document and pages can only be added but not removed.
File-based documents can also be saved and re-loaded to continue adding pages or converting to final document at a later time. For more information and examples, refer to Programming with the LEADTOOLS .NET OCR.
Use IOcrDocumentManager.CreateDocument(string, OcrCreateDocumentOptions) without the OcrCreateDocumentOptions.InMemory flag to create such documents.
IOcrDocument.IsInMemory will be false for memory-based documents. The current back-end file name can be obtained through IOcrDocument.FileName.
Typical OCR operation using the IOcrEngine involves starting up the engine, create an IOcrDocument object using the IOcrDocumentManager.CreateDocument method and adding the pages into it and perform either automatic or manual zoning. Once this is done, After the recognition data is collected using IOcrPage.Recognize, you use the various IOcrDocument.Save methods to save the document to its final format such as PDF, DOC or HTML.
In addition to the above, you can use IOcrDocument.SaveXml to save the document as XML.
This example save a page to all the formats supported by the LEADTOOLS OCR Advantage engine.
using Leadtools;using Leadtools.Codecs;using Leadtools.Forms.Ocr;using Leadtools.Forms;using Leadtools.Forms.DocumentWriters;using Leadtools.WinForms;public void OcrDocumentManagerExample(){string tifFileName1 = Path.Combine(LEAD_VARS.ImagesDir, "Ocr1.tif");string tifFileName2 = Path.Combine(LEAD_VARS.ImagesDir, "Ocr2.tif");string outputDirectory = Path.Combine(LEAD_VARS.ImagesDir, "OutputDirectory");// Create the output directoryif (Directory.Exists(outputDirectory))Directory.Delete(outputDirectory, true);Directory.CreateDirectory(outputDirectory);// Create an instance of the engineusing (IOcrEngine ocrEngine = OcrEngineManager.CreateEngine(OcrEngineType.Advantage, false)){// Start the engine using default parametersConsole.WriteLine("Starting up the engine...");ocrEngine.Startup(null, null, null, LEAD_VARS.OcrAdvantageRuntimeDir);// Create the OCR documentConsole.WriteLine("Creating the OCR document...");IOcrDocumentManager ocrDocumentManager = ocrEngine.DocumentManager;using (IOcrDocument ocrDocument = ocrDocumentManager.CreateDocument()){// Add the pages to the documentConsole.WriteLine("Adding the pages...");ocrDocument.Pages.AddPage(tifFileName1, null);ocrDocument.Pages.AddPage(tifFileName2, null);// Recognize the pages to this document. Note, we did not call AutoZone, it will explicitly be called by RecognizeConsole.WriteLine("Recognizing all the pages...");ocrDocument.Pages.Recognize(null);// Save to all the formats supported by this OCR engineArray formats = Enum.GetValues(typeof(DocumentFormat));foreach (DocumentFormat format in formats){string friendlyName = DocumentWriter.GetFormatFriendlyName(format);Console.WriteLine("Saving (using default options) to {0}...", friendlyName);// Construct the output file name (output_directory + document_format_name + . + extension)string extension = DocumentWriter.GetFormatFileExtension(format);string outputFileName = Path.Combine(outputDirectory, format.ToString() + "." + extension);// Save the documentocrDocument.Save(outputFileName, format, null);// If this is the LTD format, convert it to PDFif (format == DocumentFormat.Ltd){Console.WriteLine("Converting the LTD file to PDF...");string pdfFileName = Path.Combine(outputDirectory, format.ToString() + "_pdf.pdf");DocumentWriter docWriter = ocrEngine.DocumentWriterInstance;docWriter.Convert(outputFileName, pdfFileName, DocumentFormat.Pdf);}}// Now save to all the engine native formats (if any) supported by the enginestring[] engineFormats = ocrDocumentManager.GetSupportedEngineFormats();foreach (string engineFormat in engineFormats){string friendlyName = ocrDocumentManager.GetEngineFormatFriendlyName(engineFormat);Console.WriteLine("Saving to engine native format {0}...", friendlyName);// Construct the output file name (output_directory + "engine" + engine_format_name + . + extension)string extension = ocrDocumentManager.GetEngineFormatFileExtension(engineFormat);string outputFileName = Path.Combine(outputDirectory, "engine_" + engineFormat + "." + extension);// To use this format, set it in the IOcrDocumentManager.EngineFormat and do a normal save using DocumentFormat.User// Save the documentocrDocumentManager.EngineFormat = engineFormat;ocrDocument.Save(outputFileName, DocumentFormat.User, null);}}// Shutdown the engine// Note: calling Dispose will also automatically shutdown the engine if it has been startedConsole.WriteLine("Shutting down...");ocrEngine.Shutdown();}}static class LEAD_VARS{public const string ImagesDir = @"C:\Users\Public\Documents\LEADTOOLS Images";public const string OcrAdvantageRuntimeDir = @"C:\LEADTOOLS 19\Bin\Common\OcrAdvantageRuntime";}
Imports LeadtoolsImports Leadtools.CodecsImports Leadtools.Forms.OcrImports Leadtools.FormsImports Leadtools.Forms.DocumentWritersImports Leadtools.WinForms<TestMethod>Public Sub OcrDocumentManagerExample()Dim tifFileName1 As String = Path.Combine(LEAD_VARS.ImagesDir, "Ocr1.tif")Dim tifFileName2 As String = Path.Combine(LEAD_VARS.ImagesDir, "Ocr2.tif")Dim outputDirectory As String = LEAD_VARS.ImagesDir' Create an instance of the engineUsing ocrEngine As IOcrEngine = OcrEngineManager.CreateEngine(OcrEngineType.Advantage, False)' Start the engine using default parametersConsole.WriteLine("Starting up the engine...")ocrEngine.Startup(Nothing, Nothing, Nothing, LEAD_VARS.OcrAdvantageRuntimeDir)' Create the OCR documentConsole.WriteLine("Creating the OCR document...")Dim ocrDocumentManager As IOcrDocumentManager = ocrEngine.DocumentManagerUsing ocrDocument As IOcrDocument = ocrDocumentManager.CreateDocument()' Add the pages to the documentConsole.WriteLine("Adding the pages...")ocrDocument.Pages.AddPage(tifFileName1, Nothing)ocrDocument.Pages.AddPage(tifFileName2, Nothing)' Recognize the pages to this document. Note, we did not call AutoZone, it will explicitly be called by RecognizeConsole.WriteLine("Recognizing all the pages...")ocrDocument.Pages.Recognize(Nothing)' Save to all the formats supported by this OCR engineDim formats As Array = [Enum].GetValues(GetType(DocumentFormat))For Each format As DocumentFormat In formatsDim friendlyName As String = DocumentWriter.GetFormatFriendlyName(format)Console.WriteLine("Saving (using default options) to {0}...", friendlyName)' Construct the output file name (output_directory + document_format_name + . + extension)Dim extension As String = DocumentWriter.GetFormatFileExtension(format)Dim outputFileName As String = Path.Combine(outputDirectory, format.ToString() & "." & extension)' Save the documentocrDocument.Save(outputFileName, format, Nothing)' If this is the LTD format, convert it to PDFIf format = DocumentFormat.Ltd ThenConsole.WriteLine("Converting the LTD file to PDF...")Dim pdfFileName As String = Path.Combine(outputDirectory, format.ToString() & "_pdf.pdf")Dim docWriter As DocumentWriter = ocrEngine.DocumentWriterInstancedocWriter.Convert(outputFileName, pdfFileName, DocumentFormat.Pdf)End IfNext' Now save to all the engine native formats (if any) supported by the engineDim engineFormats As String() = ocrDocumentManager.GetSupportedEngineFormats()For Each engineFormat As String In engineFormatsDim friendlyName As String = ocrDocumentManager.GetEngineFormatFriendlyName(engineFormat)Console.WriteLine("Saving to engine native format {0}...", friendlyName)' Construct the output file name (output_directory + "engine" + engine_format_name + . + extension)Dim extension As String = ocrDocumentManager.GetEngineFormatFileExtension(engineFormat)Dim outputFileName As String = Path.Combine(outputDirectory, "engine_" & engineFormat & "." & extension)' To use this format, set it in the IOcrDocumentManager.EngineFormat and do a normal save using DocumentFormat.User' Save the documentocrDocumentManager.EngineFormat = engineFormatocrDocument.Save(outputFileName, DocumentFormat.User, Nothing)NextEnd Using' Shutdown the engine' Note: calling Dispose will also automatically shutdown the engine if it has been startedConsole.WriteLine("Shutting down...")ocrEngine.Shutdown()End UsingEnd SubPublic NotInheritable Class LEAD_VARSPublic Const ImagesDir As String = "C:\Users\Public\Documents\LEADTOOLS Images"Public Const OcrAdvantageRuntimeDir As String = "C:\LEADTOOLS 19\Bin\Common\OcrAdvantageRuntime"End Class
Leadtools.Forms.DocumentWriters.DocumentFormat
Programming with the LEADTOOLS .NET OCR
|
Products |
Support |
Feedback: IOcrDocumentManager Interface - Leadtools.Forms.Ocr |
Introduction |
Help Version 19.0.2017.6.6
|

Raster .NET | C API | C++ Class Library | JavaScript HTML5
Document .NET | C API | C++ Class Library | JavaScript HTML5
Medical .NET | C API | C++ Class Library | JavaScript HTML5
Medical Web Viewer .NET
Your email has been sent to support! Someone should be in touch! If your matter is urgent please come back into chat.
Chat Hours:
Monday - Friday, 8:30am to 6pm ET
Thank you for your feedback!
Please fill out the form again to start a new chat.
All agents are currently offline.
Chat Hours:
Monday - Friday
8:30AM - 6PM EST
To contact us please fill out this form and we will contact you via email.