| Term | Definition |
|---|---|
| Document | A set of one or more page images and the data extracted from them. |
| Document Definition | Defines how a particular document type is identified and processed: the document structure (the allowed page order, used to assemble pages into documents), document sections, the rules field data must satisfy, the locations of fields and their captions on the data form, export settings, and processing settings. |
| Document type | Documents that share characteristics and are handled uniformly within a business process. Examples: invoices, contracts, and passports. |
| Entity | A field or group of fields containing information to extract with NLP. Examples: people, companies, places, amounts, and dates. |
| Field | A document element intended for data extraction. Fields can be simple or complex. An example of a complex field is a Table field, where each cell is a separate child field. |
| NER (Named Entity Recognition) | An information-extraction task that locates and classifies named-entity mentions in unstructured text. |
| NLP (Natural Language Processing) | A subfield of artificial intelligence and computational linguistics that studies the computer analysis and synthesis of natural languages. Uses include information extraction, machine translation, chatbots, document classification, and sentiment analysis. |
| NLP model | A mechanism that determines which entities and segments to extract from text, and how. The subject area and extraction algorithm are selected during training. |
| Segment | A text fragment of one or more paragraphs that contains data to extract. A segment can also be a field to extract, for example the conditions for terminating an agreement. |
| Segmentation | The process of identifying segments. It precedes information extraction and helps with large documents by narrowing the search for entities to specific fragments. |
Using NLP to process unstructured documents
Glossary
Glossary of NLP terms used in ABBYY FlexiCapture, defining entity, field, NER, NLP model, segment, segmentation, and Document Definition.
