Form annotation for structured document extraction

Label fields, checkboxes, sections, and repeated regions in PDF forms. Give document AI models reviewed examples of where information appears and what each part means.
Forms & Structured Document Extraction example with source-data annotations and contextual photographs.

Make form structure explicit

Turn varied layouts and scanned records into consistent page-aware training annotations.
Structured PDF text, image regions, and table cells selected for extraction.

Field names and values

Mark the label and value regions of a form so models can learn their distinct roles in the page layout.
Form checkboxes labeled Checked and Unchecked.

Checkbox and choice regions

Identify checked and unchecked options using region labels and task-defined properties.
PDF form with a labeled section and repeated table rows.

Sections and repeated rows

Annotate grouped form sections and repeated table rows with a consistent layout schema.
Consent statement and signature area labeled on a PDF form.

Signature and consent fields

Mark signature areas, consent statements, and related choices as page regions with task-defined classes.

Why AI Teams Choose Unitlab

Bring document annotation, shared ontologies, expert review, and dataset delivery into one workflow so your team can focus on useful training data.
15X
Faster Document Annotation
60%
Free Up AI Engineers’ Time
5X
Lower AI Development Costs

Annotation types for forms & structured document extraction

Pair page-aware visual regions with structured labels, text values, and properties for the document task.
Structured PDF text, image regions, and table cells selected for extraction.

Form regions

Use boxes and polygons to label fields, choices, and sections on the correct PDF page.

Field values and states

Capture text values or selected-state properties according to the choices defined in your annotation ontology.

Form checkboxes labeled Checked and Unchecked.

Forms & Structured Document Extraction FAQs

What is structured form annotation?

Structured form annotation labels the positions and roles of information on a form, including field names, values, choices, and sections.

Can I annotate checkboxes?

Yes. Label checkbox regions and record their state using your project’s classes or properties. The annotation describes the visual evidence on the form.

Can I label scanned forms?

Yes. PDF pages can be annotated with visual region tools even when the underlying text is not selectable.

How do I separate a field name from its value?

Use distinct classes or a defined property scheme for the label and value regions. Keep that rule consistent across different form layouts.

Can I work with multipage forms?

Yes. Navigate the PDF’s pages while retaining the document as one work item. Saved annotations include the relevant page number.

Does the workflow require a fixed form layout?

No single layout is required. Your team can label varied forms using a shared ontology that defines the semantic roles to extract.

Which document format does Unitlab support?

The native document workflow supports PDF files. Each PDF remains one work item, with annotations attached to their original page numbers as teams navigate the document.

How are document annotations reviewed?

Reviewers inspect page regions and their labels, resolve comments, and correct or reject work through the configured workflow. Shared classes and properties keep document labeling consistent.

Does export preserve document page context?

Yes. Unitlab Unified Export Format preserves document metadata and page numbers alongside labels, properties, and supported relationships for downstream document AI pipelines.