Translated speech often carries more than words: a speaker’s identity and intonation also shape how a message sounds. Google’s AudioPaLM is designed to translate speech while retaining those vocal qualities, potentially letting someone hear a translated version in the original speaker’s voice.
One system for text and speech
AudioPaLM brings together Google’s PaLM-2 language model, introduced in May, and its AudioLM generative audio model. The combined architecture can process and generate both text and speech, supporting tasks such as speech recognition, transcription and translation.
For voice-conditioned output, the system needs a three-second audio sample. That sample is supplied as audio and a SoundStream token. If the supplied recording is shorter than three seconds, it is repeated to reach the required length.
This approach gives the model a reference for the speaker’s voice when generating translated speech. The goal is to carry elements such as speaker identity and prosody into the result, rather than turning speech into text and producing a voice that sounds unrelated to the original.
Keeping a voice recognizable across languages
AudioLM contributes the ability to generate audio with long-term consistency. According to the source, AudioPaLM can produce plausible speech continuations while maintaining speaker identity and prosody, including for speakers it did not encounter during training.
The system is also described as capable of zero-shot speech-to-text translation across many languages. That includes combinations of languages not seen during training, which could matter in practical situations where a system needs to handle speech without having been specifically prepared for every language pairing.
AudioPaLM can produce transcripts in the original language or as a translation. It can also generate speech in the source language. The source reports top results on speech translation benchmarks and competitive performance on speech recognition tasks, while saying its speech quality is expected to outperform existing solutions based on automatic and human evaluation.
Where multilingual speech could be used
The technology could support multilingual voice assistants and automated transcription services, as well as other systems that need to understand or generate written and spoken language. In these settings, a system could potentially move between listening, translating, writing and speaking within one architecture.
Video is another possible application. Google sees potential for multilingual subtitles and dubbing, including on YouTube. A dubbed version that preserves the original speaker’s voice could make translated videos feel more continuous with the source, while subtitles offer a written route to the same translated content.
These are potential uses, not a description of a finished service. The source presents AudioPaLM as a research system and points to demonstrations, while highlighting open questions that researchers still need to address.
Research questions remain
The researchers identify the properties of audio tokens as an area for further study, including how to measure and optimize them. They also call for established benchmarks and metrics for generative audio tasks. Such measures would help researchers compare systems and assess progress more consistently.
AudioPaLM brings speech generation and translation closer together, with the distinctive aim of retaining aspects of a speaker’s voice across languages. Its possible applications range from assistants and transcription to translated video, but the work described also makes clear that evaluating generative audio remains an active research need.