Skip to main content

Kyutai’s journey towards scaling emotionally expressive multilingual text-to-speech (TTS)

Learn how Kyutai built production-ready multilingual voice data for expressive TTS

1 September 2026 | United States

Black text 'kyutai' with a green forward slash on a white background.

Executive summary

As open-science AI lab Kyutai advanced multilingual text-to-speech models, it needed training data that captured more than intelligible speech. The models required studio-grade recordings with consistent pronunciation, prosody, emotion and annotation across multiple languages.

By partnering with Uber AI Solutions (UAIS), Kyutai established a full-stack speech data programme spanning voice talent, recording, transcription, annotation, metadata and quality assurance. The collaboration created a repeatable multilingual pipeline capable of producing high-quality, emotionally expressive training data for advanced voice AI systems.

Key results

500 hours

Of studio-grade speech data produced across five languages

95%+

Quality thresholds met across key QA dimensions

25%

Reduction in TTS failure rate on complex cases

Production-grade voice AI requires more than speech collection


Kyutai was pushing the boundaries of open voice and multimodal AI, including real-time conversational models and advanced speech generation.As the lab moved toward emotionally expressive text-to-speech, available datasets presented a limitation. Basic speech databases could support intelligibility, but they often lacked the emotional nuance, conversational variation, and multilingual consistency needed to produce natural-sounding voices.

Building the datasets introduced significant operational complexity. Speakers needed to maintain consistent identities and recording quality across sessions while delivering varied emotions and speaking styles. Transcripts, emotion labels, non-speech events, pronunciation, and metadata also had to remain accurate across languages and recording batches. To scale, Kyutai needed a governed production process capable of creating expressive multilingual data consistently.

Creating a unified production pipeline for multilingual voice data


Kyutai partnered with Uber AI Solutions to operationalise voice data creation as one coordinated system rather than managing recording, transcription, annotation and QA independently.

"

Uber AI Solutions is one of our most reliable data providers and annotators. They swiftly adapt to new use cases and specifications, deliver on time and with spotless quality.

"


Alexandre Défossez
Chief Exploration Officer, Kyutai

UAIS coordinated multilingual recording programmes across French, Brazilian Portuguese, European and Latin American Spanish and German. Voice talent received structured guidance designed to expand emotional and prosodic range, including empathic, sarcastic, angry, sad, crying and other high-intensity speech styles.

The workflow combined AI-assisted transcription with human verification for verbatim transcripts, emotion labels, non-speech events and other metadata. Standardised quality criteria were applied across speakers and recording batches, making QA a continuous part of production rather than a final checkpoint.

This operating model helped Kyutai maintain consistency as the programme expanded across languages, speakers and increasingly nuanced voice requirements.

Delivering a detailed dataset for expressive multilingual TTS


The engagement produced a structured source of multilingual training data tailored to Kyutai's expressive TTS requirements.

Beyond the dataset itself, the collaboration established a repeatable process for managing talent, recording quality, annotation and acceptance criteria across distributed language programmes. That foundation reduced variability before data entered model training and gave Kyutai a more consistent way to generate future speech datasets as requirements evolved.

KPI's

Acceptance Rate

Reporting frequency

Audio quality rating

95%+

With each batch submission

Pronunciation & prosody consistency

95%+

With each batch submission

Verbatim transcription accuracy

95%+

With each batch submission

Annotation accuracy (emotion & non-speech tags)

95%+

With each batch submission

Spelling & punctuation

95%+

With each batch submission

Establishing the data foundation for expressive AI voice generation


With a scalable multilingual speech data pipeline in place, Kyutai extended its work into additional languages, dialects, emotional styles and advanced conversational voice applications. For Kyutai, the partnership created the operational foundation for generating the expressive, high-quality human speech data essential to advancing increasingly natural voice AI systems.

Learn how Uber AI Solutions can partner with your enterprise to further your AI.

Testimonials reflect individual results. Individual results do not guarantee the outcome for any given customer, and customer experience will vary.