Dataset Management and Versioning for AI Teams

Manage your annotated data in one workspace. Visually inspect results, filter the examples you need, track dataset versions, and prepare reviewed releases for training.
Unitlab AI dataset management interface

Review and Filter Annotated Data

Inspect your annotated examples in context, identify issues, and focus on the data your next training run needs.

Inspect Data Before Training

Check source quality and annotation results in the relevant views. Resolve issues through review and rework before release.

Explore annotation quality assurance
A visual dataset grid with blurred, dark, corrupted, and outlier examples separated for review.

Focus on the Right Examples

Use the available data and project filters to narrow your working set. Inspect relevant subsets before review or delivery.

Explore data curation
Relevant media examples selected from a larger dataset into a focused subset.

From Reviewed Data to Training-Ready Releases

Keep the final handoff clear: review annotated examples, identify the source version, and release the output your training pipeline needs.

01

Review annotated data

Inspect the source and labels together, then resolve issues before delivery.

02

Identify the dataset version

Keep the selected source membership and its version history traceable.

03

Release and export

Freeze project annotations in a release, inspect the output, and export for training.

Dataset Version Control and Lineage

Keep source membership and annotation output traceable, from the dataset version selected for a project to the release delivered for training.

Keep Every Training Handoff Traceable

1. Preserve Source Membership

Publish a dataset version to preserve its folders and assets. Select the exact version when creating a project.

2. Inspect Version History

Review additions, removals, timestamps, and publishers to understand how the dataset changed.

3. Freeze Reviewed Annotation Output

Create a project release for delivery. Dataset versions preserve source membership; releases preserve annotation output.

Successive Unitlab dataset views showing versioned collections of annotated images.

Reuse Data with a Clear Starting Point

Create a project from a selected dataset version and work with an independent copy. Keep later annotation work separate from the source dataset.
Four images showing cars with blue overlays highlighting segmented car areas in various outdoor settings.
Three horizontal dashed arrows pointing right under the label 'Clone Dataset'.Three vertical dashed arrows pointing downward under the text 'Clone Dataset'.
Dashboard showing progress of car image segmentation with four segmented car photos and 20% label and review progress.

Connected data. One shared context.

Explore Multimodal Annotation

Object detection, segmentation, and keypoints.

Explore Image Annotation

Object tracking and frame-accurate labels.

Explore Video Annotation

DICOM and volumetric medical imaging.

Explore Medical Annotation

Whole-slide images, tissue, and nuclei.

Explore Pathology Annotation

Satellite and aerial imagery.

Explore Geospatial Annotation

Entities, relationships, and text classification.

Explore Text Annotation

PDFs, OCR fields, and document layouts.

Explore Document Annotation

Text spans in their original webpage context.

Explore HTML Annotation

Speech, speakers, and sound events.

Explore Audio Annotation

Time-series intervals and sampled events.

Explore Sensor Annotation

Record, cell, and field-level annotation.

Explore Tabular Annotation

Dataset Management FAQs

What is dataset management for AI training data?

Dataset management helps AI teams turn already annotated data into organized, traceable training inputs. In Unitlab, teams can visually inspect examples, filter relevant data, manage versions, and prepare reviewed outputs for training. It connects the results of data annotation with a repeatable delivery workflow.

How does dataset management differ from data curation?

Data curation focuses on finding and selecting relevant source data. Dataset management focuses on organizing, visually checking, filtering, and versioning annotated data for downstream use. Together, they connect deliberate data selection with traceable training-data delivery.

Can teams visually inspect and filter annotated datasets?

Yes. Inspect the underlying data and annotations, then use available filters and properties to narrow the examples you need to check. Resolve labeling issues through quality assurance, review, and rework before preparing a training release.

What is the difference between a dataset version and an annotation release?

A dataset version preserves the source-data membership of a dataset at a specific point. An annotation release freezes a project's annotation output for delivery. Use dataset versions to identify your input data and project releases to identify the labeled output used for training.

Can I reuse a dataset version in another annotation project?

Yes. A published dataset version can provide source data for another project. The project works with its own copy, so later annotation changes do not modify the source dataset. See dataset versions and history for the workflow.

Which data modalities can I manage?

Unitlab supports workflows for images, video, audio, text, documents and PDFs, HTML webpages, medical imaging, whole-slide pathology, geospatial imagery, sensor and time-series data, and tabular data. Explore multimodal annotation when related data needs shared context, and choose the appropriate output format for each training workflow.

How do I prepare reviewed data for model training?

Inspect the examples and resolve outstanding issues through your annotation quality workflow. Record the source dataset version, create a project release of the reviewed annotations, and export in a supported format suited to your training pipeline. This makes the input data and delivered labels easier to trace.

Can dataset workflows connect to our existing tools?

Use the API documentation, Python SDK guide, and CLI documentation to identify supported integration steps. Keep dataset-version and release identifiers alongside training runs so your team can track which data and annotations were used.

Want to evaluate your dataset workflow? Book a platform demo.