How VALL-E Uses Three Seconds to Recreate a Voice

Microsoft’s VALL-E turns text into speech using a three-second acoustic prompt to guide the voice. Trained on 60,000 hours of English speech from 7,000 speakers, it can also carry features such as emotion and recording conditions into generated audio.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

VALL-E’s voice imitation raises some potential for impersonation, while the story mainly describes a speech synthesis model.

How VALL-E Uses Three Seconds to Recreate a Voice

Microsoft’s VALL-E is a text-to-speech model that uses a short audio sample to shape how written words sound. In the researchers’ tests, a three-second acoustic prompt helped the system generate speech in the indicated voice while retaining some qualities of the recording.

A different route from text to speech

Many text-to-speech systems convert phonemes into spectrograms, process those representations in a neural network, and then produce audio waveforms. VALL-E takes another route: it represents the text and acoustic prompt as tokens, then uses an audio codec decoder to turn those tokens into waveforms.

The approach draws on methods associated with large language models. The researchers describe training a model with large and varied data as an alternative to building a complex network specifically for the task. VALL-E’s design applies that broad-data approach to speech synthesis.

Training on a large speech collection

The model was trained on 60,000 hours of English speech from 7,000 speakers. The Microsoft team says that amount is more than 100 times the data previously used in the field. The training material came from the LibriLight dataset, which the researchers transcribed with AI.

That scale matters in the context of a challenge faced by speech synthesis. Models trained on high-quality recordings can perform poorly with lower-quality audio, and generated speech quality can drop for speakers absent from the training set. Earlier approaches to these zero-shot cases have used speaker adaptation or speaker encoding, which can require fine-tuning or pre-designed features.

VALL-E instead uses a brief sample of the target voice as an acoustic prompt. Paired with a text prompt, that sample gives the model information about how to produce the requested speech.

What the acoustic prompt carries over

The researchers’ examples show the system generating text in the voice represented by the audio prompt. They also report that aspects of the recording can carry into the output. A phone recording’s noise, for instance, may remain audible in the synthesized continuation.

The same idea applies to other qualities. VALL-E can preserve pitch influenced by emotion, such as an angry speaker’s pitch, and reproduce reverberation when it is present in the prompt. The researchers say a baseline system produced clean speech in the reverberation comparison, while VALL-E could retain the acoustic environment.

They attribute that consistency to training on a large dataset with a wider range of acoustic conditions. Their interpretation is that VALL-E learned to continue those conditions rather than assuming every recording should sound clean.

Capabilities and limits

The researchers say they did not explicitly train VALL-E to integrate emotions into generated voices. They describe its ability to preserve emotion and acoustic environment as an emergent capability, and say the model can learn in context in a way they compare with large language models.

These results point to a practical strength: a short voice sample can provide more than a speaker identity. It can also guide qualities of the resulting audio, from recording texture to emotional pitch. That makes the model’s prompt an important part of how each output sounds.

The available material has limits. The researchers shared audio examples on GitHub, but the code is not available. The reported demonstrations therefore offer examples of VALL-E’s behavior, while leaving readers without code to reproduce the system themselves.