Data Platform
TTS & STT · browser-based capture

High-quality speech data, collected right in the browser.

Import prompts, record clean takes with voice-call DSP disabled, and let the pipeline resample, normalize, and quality-gate every clip — then export training-ready datasets for text-to-speech and speech recognition.

Everything for a clean speech corpus

From the first prompt to the exported dataset bundle.

In-browser recording

A no-reload capture loop with live level metering, device selection, and voice-call DSP (echo cancellation, noise suppression, AGC) turned off for faithful audio.

TTS & STT pipelines

Per-task pipelines — 22.05/24 kHz studio takes for TTS, 16 kHz natural speech for ASR — each with its own resampling and quality gates.

Automatic quality gates

Every clip is resampled, loudness-normalized, silence-trimmed, and checked for duration and speech presence before it counts as passed.

One-click dataset export

Bundle passed clips into a manifest + meta + WAVs, filter by task / style / language, and ship straight to S3 — reproducible and versioned.

Three steps to a dataset

Built for teams of contributors collecting in parallel.

  1. 1

    Import prompts

    Upload a CSV or TSV of texts and tag them by task, label, language, and style.

  2. 2

    Record

    Contributors work through their queue — read the prompt, capture, auto-advance. Each take is uploaded and processed in the background.

  3. 3

    Export

    Review passed clips on the dashboard, then export a versioned bundle locally or to S3.

Ready to start collecting?

Sign in with your contributor or admin account to open your recording queue.

Sign in