





Select one contiguous, non-overlapping text range for each product field. Assign its class while preserving the exact visible source characters.
Connect a specification name to its value within the same HTML page. Add required properties to describe field roles or human-entered normalized values.

It labels the visible product fields that a model should extract from a webpage, such as a name, brand, model, price or specification value. Unitlab keeps those labels connected to the uploaded HTML snapshot for training and evaluation.
Define entity classes for the visible fields your dataset requires. Common examples include product names, brands, model identifiers, prices, availability, materials, dimensions and other stated specifications.
Yes. Use different classes or a defined price-role property, and inspect the surrounding offer text. Preserve the actual source value rather than assuming every amount on the page describes the main product.
Yes. Annotate each visible name and value as a separate text entity, then create an ontology-defined relation between them within the same HTML page. The relation records the connection your extraction task needs.
Yes, when your ontology includes suitable properties. Annotators can enter reviewed normalized values while retaining the selected source text. This workflow does not automatically infer or normalize product attributes.
The HTML workflow uses uploaded .html or .htm snapshots. Annotators review the rendered saved content; it is not a live website crawler or a browser session that follows links and changes the page.
This solution prepares and reviews human-labeled extraction examples. It does not imply automatic crawling, product matching or automatic field extraction. Your downstream model can use the reviewed labels according to its training requirements.
Review the selected text in its rendered context, including nearby titles, offer details and recommendations. Contextual comments and the configured review workflow help return incorrect occurrences or ambiguous field roles for correction.
HTML annotations support JSONL and Unitlab Unified Export Format. Exports preserve supported source identity, text-anchor information, entity labels, properties, relations and item properties for downstream dataset preparation.