Meta’s Speech Models Reach 1,100 Languages

Meta’s Massively Multilingual Speech project offers models for speech recognition and text-to-speech across 1,100 languages. The work combines New Testament recordings with broader speech training, while Meta cautions that transcription errors can still produce incorrect or offensive statements.

WTF Index IDIOCRACY
◄ Terminator 0 Idiocracy 1 ►

The project broadens speech access, though transcription errors can produce incorrect or offensive statements.

Meta’s Speech Models Reach 1,100 Languages

Speech technology often works across far fewer languages than people speak. Meta’s Massively Multilingual Speech project aims to widen that reach with open-source AI models that convert speech to text and text to speech in 1,100 languages, and can identify more than 4,000.

Building speech tools across languages

The models are based on Meta’s wav2vec and were trained using audio and text resources at different scales. A curated dataset includes examples for 1,100 languages. Meta also assembled an uncurated dataset covering nearly 4,000 languages, including some spoken by only a few hundred people and for which speech technology does not yet exist.

That broader coverage is the project’s central contribution. Meta says MMS covers ten times more languages than previous models. Recognition, identification and speech generation are distinct tasks, but the project’s models address all three across languages that are rarely represented in existing speech tools.

New Testament recordings provide training material

A key part of the dataset is the New Testament. Meta’s collection includes readings in more than 1,107 languages, averaging 32 hours per language. The recordings were paired with matching passages from the Internet, giving the researchers audio alongside text.

The project also used 3,809 unlabeled audio files. These were New Testament readings without additional language information. Together, the labeled and unlabeled recordings helped create training material across a wide range of languages.

Meta says 32 hours per language is not enough to train a reliable speech recognition system by itself. To address that limit, it used wave2vec 2.0 to pre-train MMS models with more than 500,000 hours of speech in more than 1,400 languages. The models were then fine-tuned to understand or identify numerous languages.

Coverage does not remove accuracy limits

Benchmarks reported in the article found that performance remained nearly constant as the number of training languages grew. The error rate decreased minimally by 0.4 percentage points with increasing training, according to the report.

Meta says MMS has a significantly lower error rate than OpenAI’s Whisper, which was not explicitly optimized for extensive multilingualism. The article notes that a comparison in English alone would be more informative, and that first testers on Twitter reported Whisper performed better there. Those observations point to a distinction between broad language coverage and performance in a particular language or task.

Meta also cautions that the model can transcribe words or phrases incorrectly. Such errors may lead to statements that are wrong or offensive, so the breadth of languages supported does not mean every result is reliable.

The article reports that the voices in the dataset are predominantly male. Meta says this does not negatively affect understanding or generation of female voices. It also says the model does not tend to generate overly religious speech, attributing this to Connectionist Temporal Classification, an approach focused more on speech patterns and sequences than on word content and meaning.

Open models and a longer-term goal

Meta’s stated long-term goal is a single language model that supports as many languages as possible, with the aim of preserving endangered languages. The company says future models may add more languages and dialects. It also describes possible uses in VR and AR technologies and messaging, where people could access information or use devices in their preferred language.

Meta says a future model could be trained across tasks such as speech recognition, speech synthesis and speech identification, potentially improving overall performance. For now, the released resources include the code, pre-trained MMS models with 300 million and one billion parameters, and refined versions for speech recognition, identification and text-to-speech. Meta has made these available as open-source models on Github.

The project shows how a large collection of recordings can extend speech tools to languages that have had little or no technology support. Its reach is broad, but Meta’s own caution about transcription mistakes remains important: language availability is a step toward access, while dependable results still require care.