Skip to main content

Kyutai's path to scaling emotionally expressive multilingual text-to-speech (TTS)

Learn how Kyutai built production-ready multilingual voice data for expressive TTS

September 1, 2026 | United States

Black text 'kyutai' with a green forward slash on a white background.

Executive summary

As open-science AI lab Kyutai advanced multilingual text-to-speech models, it needed training data that captured more than intelligible speech. The models required studio-grade recordings with consistent pronunciation, prosody, emotion, and annotation across multiple languages.

By partnering with Uber AI Solutions (UAIS), Kyutai established a full-stack speech data program spanning voice talent, recording, transcription, annotation, metadata, and quality assurance. The collaboration created a repeatable multilingual pipeline capable of producing high-quality, emotionally expressive training data for advanced voice AI systems.

Key results

500 hours

Of studio-grade speech data produced across five languages

95%+

Quality thresholds achieved across key QA dimensions

25%

Reduction in TTS failure rate on complex cases

Production-grade voice AI requires more than speech collection


Kyutai was pushing the boundaries of open voice and multimodal AI, including real-time conversational models and advanced speech generation.As the lab moved toward emotionally expressive text-to-speech, available datasets presented a limitation. Basic speech databases could support intelligibility, but they often lacked the emotional nuance, conversational variation, and multilingual consistency needed to produce natural-sounding voices.

Building the datasets introduced significant operational complexity. Speakers needed to maintain consistent identities and recording quality across sessions while delivering varied emotions and speaking styles. Transcripts, emotion labels, non-speech events, pronunciation, and metadata also had to remain accurate across languages and recording batches. To scale, Kyutai needed a governed production process capable of creating expressive multilingual data consistently.

Creating a unified production pipeline for multilingual voice data


Kyutai partnered with Uber AI Solutions to operationalize voice data creation as one coordinated system rather than managing recording, transcription, annotation, and QA independently.

Uber AI Solutions is one of our most reliable data providers and annotators. They swiftly adapt to new use cases and specifications, deliver on-time, and with spotless quality.


Alexandre Défossez
Chief Exploration Officer, Kyutai

UAIS coordinated multilingual recording programs across French, Brazilian Portuguese, European and Latin American Spanish, and German. Voice talent received structured guidance designed to expand emotional and prosodic range, including empathic, sarcastic, angry, sad, crying, and other high-intensity speech styles.

The workflow combined AI-assisted transcription with human verification for verbatim transcripts, emotion labels, non-speech events, and other metadata. Standardized quality criteria were applied across speakers and recording batches, making QA a continuous part of production rather than a final checkpoint.

This operating model helped Kyutai maintain consistency as the program expanded across languages, speakers, and increasingly nuanced voice requirements.

Delivering a detailed dataset for expressive multilingual TTS


The engagement produced a structured source of multilingual training data tailored to Kyutai's expressive TTS requirements.

Beyond the dataset itself, the collaboration established a repeatable process for managing talent, recording quality, annotation, and acceptance criteria across distributed language programs. That foundation reduced variability before data entered model training and gave Kyutai a more consistent way to generate future speech datasets as requirements evolved.

KPI’s

Acceptance Rate

Reporting Frequency

Audio quality score

95%+

With each batch submission

Pronunciation & prosody consistency

95%+

With each batch submission

Verbatim transcription accuracy

95%+

With each batch submission

Annotation accuracy (emotion & non speech tags)

95%+

With each batch submission

Spelling & punctuation

95%+

With each batch submission

Establishing the data foundation for expressive AI voice generation


With a scalable multilingual speech data pipeline in place, Kyutai extended its work into additional languages, dialects, emotional styles, and advanced conversational voice applications. For Kyutai, the partnership created the operational foundation for generating the expressive, high-quality human speech data essential to advancing increasingly natural voice AI systems.

Learn how Uber AI Solutions can partner with your enterprise to advance your AI.

Testimonials reflect individual results. Individual results are not a guarantee of outcome for any given customer, and customer experience will vary.