Kyutai's path to scaling emotionally expressive multilingual text-to-speech (TTS)
Learn how Kyutai built production-ready multilingual voice data for expressive TTS
September 1, 2026 | United States
Executive summary
As open-science AI lab Kyutai advanced multilingual text-to-speech models, it needed training data that captured more than intelligible speech. The models required studio-grade recordings with consistent pronunciation, prosody, emotion, and annotation across multiple languages.
By partnering with Uber AI Solutions (UAIS), Kyutai established a full-stack speech data program spanning voice talent, recording, transcription, annotation, metadata, and quality assurance. The collaboration created a repeatable multilingual pipeline capable of producing high-quality, emotionally expressive training data for advanced voice AI systems.
Key results
500 hours
Of studio-grade speech data produced across five languages
95%+
Quality thresholds achieved across key QA dimensions
25%
Reduction in TTS failure rate on complex cases
Production-grade voice AI requires more than speech collection
Kyutai was pushing the boundaries of open voice and multimodal AI, including real-time conversational models and advanced speech generation.As the lab moved toward emotionally expressive text-to-speech, available datasets presented a limitation. Basic speech databases could support intelligibility, but they often lacked the emotional nuance, conversational variation, and multilingual consistency needed to produce natural-sounding voices.
Building the datasets introduced significant operational complexity. Speakers needed to maintain consistent identities and recording quality across sessions while delivering varied emotions and speaking styles. Transcripts, emotion labels, non-speech events, pronunciation, and metadata also had to remain accurate across languages and recording batches. To scale, Kyutai needed a governed production process capable of creating expressive multilingual data consistently.
Creating a unified production pipeline for multilingual voice data
Kyutai partnered with Uber AI Solutions to operationalize voice data creation as one coordinated system rather than managing recording, transcription, annotation, and QA independently.
“
Uber AI Solutions is one of our most reliable data providers and annotators. They swiftly adapt to new use cases and specifications, deliver on-time, and with spotless quality.
”
Alexandre Défossez
Chief Exploration Officer, Kyutai
UAIS coordinated multilingual recording programs across French, Brazilian Portuguese, European and Latin American Spanish, and German. Voice talent received structured guidance designed to expand emotional and prosodic range, including empathic, sarcastic, angry, sad, crying, and other high-intensity speech styles.
The workflow combined AI-assisted transcription with human verification for verbatim transcripts, emotion labels, non-speech events, and other metadata. Standardized quality criteria were applied across speakers and recording batches, making QA a continuous part of production rather than a final checkpoint.
This operating model helped Kyutai maintain consistency as the program expanded across languages, speakers, and increasingly nuanced voice requirements.
Delivering a detailed dataset for expressive multilingual TTS
The engagement produced a structured source of multilingual training data tailored to Kyutai's expressive TTS requirements.
Beyond the dataset itself, the collaboration established a repeatable process for managing talent, recording quality, annotation, and acceptance criteria across distributed language programs. That foundation reduced variability before data entered model training and gave Kyutai a more consistent way to generate future speech datasets as requirements evolved.
KPI’s | Acceptance Rate | Reporting Frequency |
|---|---|---|
Audio quality score | 95%+ | With each batch submission |
Pronunciation & prosody consistency | 95%+ | With each batch submission |
Verbatim transcription accuracy | 95%+ | With each batch submission |
Annotation accuracy (emotion & non speech tags) | 95%+ | With each batch submission |
Spelling & punctuation | 95%+ | With each batch submission |
Establishing the data foundation for expressive AI voice generation
With a scalable multilingual speech data pipeline in place, Kyutai extended its work into additional languages, dialects, emotional styles, and advanced conversational voice applications. For Kyutai, the partnership created the operational foundation for generating the expressive, high-quality human speech data essential to advancing increasingly natural voice AI systems.
Learn how Uber AI Solutions can partner with your enterprise to advance your AI.
Testimonials reflect individual results. Individual results are not a guarantee of outcome for any given customer, and customer experience will vary.
Solutions
Industries
Resources
Explore