





Define a speech segment on the waveform and enter its transcript as an annotation property. Refine the start and end against the actual recording.
Use structured properties for language, acoustic condition, or transcription uncertainty. Keep those labels grounded in the source audio and the dataset policy.

It is the preparation of audio segments with the words spoken in each segment and any required context labels. Unitlab supports time-based audio annotation, transcripts, structured properties, and review workflows.
Yes. Transcript text can be attached to a timed speech annotation. Annotators can adjust the interval and review the corresponding words against the source recording.
Use an agreed convention for unintelligible or uncertain speech and record the relevant quality properties. Reviewers should inspect the audio before accepting the transcript.
Yes. Define language properties and use appropriate transcription guidelines for the languages in your dataset. Assign reviewers with the expertise needed to validate those examples.
Transcription records what was said. Diarization records which speaker label applies to each time interval. A dataset can include both when the downstream speech task requires them.
Define shared classes, structured properties, and clear labeling instructions before work starts. Use representative examples and contextual review to resolve disagreements in the speech recognition dataset.
Yes. Route audio annotation through Review and Rework stages. Reviewers can inspect the source data, correct labels, and send an item back when more work is needed.
Yes. Use dataset search, metadata, tags, and available filters to select relevant audio assets. Keep representative conditions and difficult examples visible in the preparation workflow.
Export reviewed audio labels in supported Audio JSON, RTTM, or UUEF formats according to the annotation task. Dataset versions help teams identify which prepared examples belong to a release.