Meta has introduced SeamlessM4T, an AI model designed to translate both spoken and written language. It can take speech or text as input and return translated text or speech, aiming to make communication across languages easier.
One model for speech and text
SeamlessM4T handles several tasks that are often treated separately. It can recognize speech and turn it into text, translate spoken audio into text in another language, or produce translated speech from spoken input. It can also translate text into text or generate spoken output from text.
Meta says the text translation functions cover nearly 100 languages. The speech output functions support about 36 languages. That distinction matters: the headline figure of up to 100 languages does not mean every feature works with speech in every one of those languages.
The model works with both audio and text, which lets it support different kinds of translation within one system. Meta describes that single-system approach as a way to reduce errors and make the process more efficient than linking multiple models together.
Training data behind the system
Training such a model requires examples that connect speech and text across languages. Meta’s researchers describe building SeamlessAlign, a multimodal collection of automatically aligned speech translations. The research paper says the corpus contained more than 470,000 hours, then describes a filtered subset with human-labeled and pseudo-labeled data totaling 406,000 hours.
The announcement also describes SeamlessAlign as totaling 270,000 hours of mined speech and text alignments. Those figures refer to the dataset as presented in different parts of the source material; the announcement does not explain the difference between them.
For text, Meta says it used the same dataset deployed in NLLB. The source describes that material as sentences from Wikipedia, news sources, scripted speeches and other sources, translated by professional human translators.
The speech data came from 4 million hours of raw audio from a publicly available repository of crawled web data. The research paper says 1 million hours were in English. Meta did not specify which repository it used or where the audio clips came from, leaving the material’s provenance unclear in the announcement.
A release for research
Meta is releasing SeamlessM4T under a CC BY-NC 4.0 research license. That license allows developers to build on the work. The company is also releasing SeamlessAlign, which could give other researchers a resource for training translation systems.
Making a model and dataset available does not, by itself, answer questions about how well the system performs across all its supported languages or situations. The available description establishes the range of tasks and language coverage Meta claims, while leaving the source of some audio data unspecified.
Part of a wider translation effort
Machine-learning translation tools predate SeamlessM4T. Google Translate has used machine-learning techniques since 2006, and large language models such as GPT-4 can translate between languages. The source also points to OpenAI’s Whisper, a speech-to-text translation model that can recognize speech in audio and translate it to text.
SeamlessM4T extends this direction by bringing speech and text translation into one model and covering many languages. Meta compares the ambition to the fictional Babel Fish from Douglas Adams’ The Hitchhiker’s Guide to the Galaxy, while acknowledging that universal translation remains a challenge because existing speech systems cover only a small fraction of the world’s languages.
For people using translation tools, the practical promise is more ways to move between spoken and written communication. For researchers, the model and aligned data offer a starting point for further work. The size of the language list is only one measure of usefulness; the model’s support for different inputs and outputs is another.