Import prompts, record clean takes with voice-call DSP disabled, and let the pipeline resample, normalize, and quality-gate every clip — then export training-ready datasets for text-to-speech and speech recognition.
From the first prompt to the exported dataset bundle.
A no-reload capture loop with live level metering, device selection, and voice-call DSP (echo cancellation, noise suppression, AGC) turned off for faithful audio.
Per-task pipelines — 22.05/24 kHz studio takes for TTS, 16 kHz natural speech for ASR — each with its own resampling and quality gates.
Every clip is resampled, loudness-normalized, silence-trimmed, and checked for duration and speech presence before it counts as passed.
Bundle passed clips into a manifest + meta + WAVs, filter by task / style / language, and ship straight to S3 — reproducible and versioned.
Built for teams of contributors collecting in parallel.
Upload a CSV or TSV of texts and tag them by task, label, language, and style.
Contributors work through their queue — read the prompt, capture, auto-advance. Each take is uploaded and processed in the background.
Review passed clips on the dashboard, then export a versioned bundle locally or to S3.
Sign in with your contributor or admin account to open your recording queue.
Sign in