Skip to main content
A document includes fields to be filled in by hand or by machine. Documents may have one or more pages. Documents can be divided into fixed and semi-structured documents.

Fixed documents

In fixed documents, the identical fields have exactly the same locations on all the documents in a batch. Fixed documents can be processed by document processing applications which read information from the data fields and export it into databases, document management systems, or archiving applications. Data are captured from fixed documents using a Document Definition, which describes the locations of the fields and the type of information they may contain. The same Document Definition is used to capture data from all the documents in a given batch. It tells the document processing application where to look for specific data on a document and how to make sure that the data have been captured correctly.

Semi-structured documents

In semi-structured documents, the locations of identical data fields vary from one document to another. Additionally, not all fields may be present on all of the documents in a batch (for example, some documents may contain a signature field while others may not). Various payment documents are good examples of semi-structured documents. Letters, registration forms, and legal documents are other good examples of semi-structured documents. Documents of the same type will have similar structures but there may still be discrepancies among their fields. For example, letters will contain the name and address of the sender at the top of the page, and legal documents will contain the names of the parties, their details, and the effective date. Since the exact location of fields on semi-structured documents is not known in advance, data cannot be captured from such documents using a Document Definition. This means that traditional data capture systems cannot extract data from such documents.

How FlexiLayouts capture data from semi-structured documents

ABBYY FlexiLayout Studio allows you to formally describe unstructured documents and provide the data capture application with a search algorithm, enabling it to find data fields and extract information from these fields. A formal description relies on mutual relationships among the fields on an unstructured document and the nature of the data within the fields. You can test the descriptions you create on document images to make sure that information can be reliably extracted. Formalized descriptions created with ABBYY FlexiLayout Studio are called FlexiLayouts. To start capturing data from unstructured documents using a FlexiLayout, you must export it into a data capture application such as ABBYY FlexiCapture. ABBYY FlexiCapture technology offers a wide range of data capture capabilities, enabling you to process practically any type of document.