
Text Span Annotation
Select product names, job titles, prices, and other visible text directly on a rendered HTML page. Keep each label anchored to the exact words and their surrounding context.
Turn saved web pages into reliable AI training data. Label text in its original layout, connect entities, and review annotations with shared ontologies and quality workflows.
Annotate rendered pages with precise text spans, contextual relations, reusable ontologies, and traceable human review.

Select product names, job titles, prices, and other visible text directly on a rendered HTML page. Keep each label anchored to the exact words and their surrounding context.

Keep headings, tables, images, and nearby text visible while annotating. Switch between fixed-width, fit, and responsive views to inspect each saved page comfortably.

Connect labeled entities within a page, such as a product and its brand or a job and its employer. Make the relationships your extraction models need explicit.

Define reusable entity classes, nested properties, page-level fields, and allowed relationships. Keep teams aligned on the same labeling rules across HTML datasets.

Compare annotators’ labels on the same HTML page. Inspect highlighted differences and resolve disagreements against shared guidelines.

Leave feedback beside the text that needs attention. Reviewers can inspect the original page, request corrections, and keep quality decisions connected to each annotation.
Give AI teams one place to label HTML pages, review ambiguous cases, and maintain reusable training datasets.
Capture the words, relationships, and contextual fields your models need while keeping the original page visible.
Label exact words or passages as names, prices, titles, dates, descriptions, and other entities defined by your task.
Connect annotated text entities to describe relationships such as product-to-brand or job-to-employer.
Add structured attributes to labeled entities, including category, status, and other details required by your ontology.
Classify the complete HTML page with fields such as content type, topic, language, relevance, or review status.
Organize saved pages, select relevant examples, and move annotations through review into a versioned training-data release.

Bring uploaded HTML snapshots into a shared data space with page previews and source metadata. Keep source pages organized as your dataset grows.

Find saved pages by name, tags, folders, and metadata. Narrow the source collection to the examples needed for your annotation task.

Group selected pages by source, category, or task. Prepare a consistent set of examples before handing the collection to annotators.

Assign HTML pages, apply shared instructions, and route work through annotation, review, and rework. Keep reviewers and annotators connected around the same source content.

Use dataset versions to track source pages. After review, create a project annotation release to freeze the labels and export them with their source context. Keep each training-data handoff traceable.
Build reviewed training data for product information, job postings, and useful web content, with labels grounded in the pages where they appear.

Label product names, brands, prices, and specifications in saved product pages. Prepare consistent examples for catalog extraction models.
Explore product-page extraction
Identify job titles, employers, locations, skills, and requirements in rendered job postings. Build structured training data for recruitment search and matching.
Explore job-posting extraction
Separate useful article content from surrounding page text. Label titles, authors, dates, and passages for cleaner web-content datasets.
Explore web-content extractionAnswers about labeling rendered HTML, preserving page context, defining ontologies, reviewing annotations, and preparing web training datasets.
Talk with the Unitlab teamHTML annotation labels information in saved web pages to create training and evaluation data. In Unitlab, annotators work with the rendered page, select visible text, connect entities, and add structured properties while retaining the surrounding layout.
Web-content extractionYou can label visible text spans such as product names, job titles, prices, locations, authors, dates, and article passages. Entity relations connect labeled spans, class properties describe each entity, and item properties classify the page as a whole.
Product-page extractionYes. Unitlab displays an uploaded HTML snapshot so annotators can interpret text alongside headings, tables, images, and nearby content. Fixed-width, fit, and responsive viewing modes help teams inspect the same saved source at a comfortable size.
Job-posting extractionA shared ontology defines the entity classes, properties, and relations used across the project. Structured choices, required fields, and reusable rules help annotators apply the same interpretation to similar content across different page layouts.
Ontology documentationReviewers can inspect labels in the original page, leave anchored comments, and return work for correction. Consensus comparison helps teams examine differences between annotators and resolve ambiguous spans or relationships using shared guidelines.
Annotation quality assuranceYes. Use the data library, folders, tags, metadata, and filters to organize uploaded HTML pages and select the examples needed for a task. Curated collections and dataset versions keep the chosen source set traceable through annotation and review.
Data curationKeep the selected source pages traceable with a dataset version. After review, create a project annotation release to freeze the labels, then export JSONL or Unitlab’s native UUEF format. HTML exports preserve labeled text, entity properties, relationships, snapshot identity, and text anchors.
Export documentationTurn saved web pages into reviewed labels, meaningful entity relationships, and versioned datasets. Give your extraction models consistent training data with the original page context intact.