Meta Builds Its Own AI Chips as Research Scales Up

Meta is developing its MTIA chip family to make neural network inference faster and less costly, while expanding its RSC supercomputer for AI research. The system has reached its second stage, and Meta says it can use data from production systems for training.

WTF Index TERMINATOR
◄ Terminator 1 Idiocracy 0 ►

The story mildly leans toward greater AI capacity through expanded inference hardware and research computing, without describing clear harm or loss of human capability.

Meta Builds Its Own AI Chips as Research Scales Up

Meta is working on two parts of its AI infrastructure: custom chips for running neural networks and a larger supercomputer for research. The plans show how the company is building capacity both to execute AI models and to train them.

Custom hardware for running AI models

The Meta Training and Inference Accelerator, or MTIA, is a family of chips intended to accelerate inference: the processing involved when a neural network runs. Meta says the chips are designed to make that work faster and cheaper. The company expects the MTIA to be in use by 2025, while its data centers still rely on Nvidia graphics cards.

MTIA is an application-specific integrated circuit, or ASIC. Like Google's Tensor Processing Units, it is optimized for operations used in neural networks, including matrix multiplication and activation functions. Meta says its chip can handle low and medium-complexity AI models better than a GPU.

The focus on inference gives the chip a specific role. It is meant to run models, rather than serve as a general-purpose replacement for all the computing hardware in Meta's data centers. That targeted design could help the company handle some AI workloads with equipment built around those tasks.

The RSC expands Meta's research capacity

Meta's Research SuperCluster, or RSC, is the other major part of this effort. The company unveiled it in January 2022, saying it would lay groundwork for the Metaverse. Meta said the completed system was intended to be the fastest supercomputer specializing in AI calculations. Construction of the infrastructure began in 2020.

The RSC has now reached its second stage. According to Meta, it includes 2,000 Nvidia DGX A100 and 16,000 Nvidia A100 GPUs, with peak performance of five exaflops. The company plans to use it for AI research across several areas, including generative AI.

That scale supports a range of research work, from developing models to exploring generative AI. The supercomputer's role complements MTIA: the RSC provides a large system for research and training, while the new chip family targets the execution of neural networks.

Production data adds a distinctive capability

Meta says the RSC can use data from its production systems for AI training. Until now, the company has relied primarily on open-source and publicly available datasets, despite having a large collection of its own data. This capability gives Meta another potential input for its research.

The source does not detail what production data will be used or how it will shape particular projects. It does identify the ability to draw on those systems as a feature of the RSC. Together with the computing resources, that gives the infrastructure a role beyond simply supplying more processing power.

Research results and the wider chip landscape

The RSC has already been used to train Meta's LLaMA language model. Meta says training the largest LLaMA model took 21 days on 2,048 Nvidia A100 GPUs. The model became part of the open-source language model movement after it was partly leaked and partly published.

Meta's move toward custom chips sits within a broader industry interest in specialized AI hardware. Amazon offers access to its Trainium and Inferentia chips for training and execution in the cloud. Microsoft is said to be working with AMD on AI chips.

For Meta, the combination of MTIA and the RSC addresses different stages of AI work. One is designed to make model execution more efficient; the other gives researchers a large platform for training and experimentation. The company's infrastructure strategy therefore depends on both purpose-built hardware and substantial computing capacity.