Amazon’s Trainium2 and Graviton4 target two AI chip needs

Amazon introduced Trainium2 for training AI models and Graviton4 for running them in the AWS cloud. The company says Trainium2 can scale to clusters of 100,000 chips, while Graviton4 offers performance and memory improvements over Graviton3.

WTF Index TERMINATOR
◄ Terminator 1 Idiocracy 0 ►

The story describes scaling AI training and inference capacity, a mild increase in AI capability without a clear harm or dependency angle.

Amazon’s Trainium2 and Graviton4 target two AI chip needs

Amazon has introduced two new chips for different stages of AI work: Trainium2 is built to train models, while Graviton4 is intended to run trained models. The announcements come as demand for generative AI adds pressure to the supply of GPUs, which are commonly used for both jobs.

Trainium2 is built for large training workloads

AWS says Trainium2 can deliver up to four times the performance and twice the energy efficiency of the first-generation Trainium. Amazon unveiled that earlier chip in December 2020.

The new processor is planned for EC Trn2 instances, grouped in clusters of 16 chips in the AWS cloud. Amazon says those deployments can scale to 100,000 chips through its EC2 UltraCluster product. The company estimates that a cluster of that size can train a 300-billion-parameter large language model in weeks instead of months.

Parameters are the elements a model learns from training data. They help define what the model can do, such as generate text or code. For context, the model size Amazon describes is about 1.75 times that of OpenAI’s GPT-3, which came before GPT-4.

Amazon also says 100,000 Trainium chips would provide 65 exaflops of compute. Exaflops and teraflops describe how many computing operations can be performed per second. The article notes that dividing the company’s total figure across the chips gives a rough per-chip estimate, but that simple calculation may not capture all the factors involved.

AWS has not set a precise Trainium2 launch date

Amazon said Trainium2 instances would arrive “sometime next year,” without providing a more specific date. That leaves customers without a clear schedule for when they can try the new training hardware in AWS.

The chip announcement reflects the broader push by large technology companies to develop processors tailored to their own AI workloads. Custom chips can give cloud providers another way to serve customers as demand for AI training grows and GPUs remain in short supply.

AWS compute and networking VP David Brown described silicon as central to customer workloads. He said Trainium2 is intended to help customers train machine-learning models faster, at lower cost, and with better energy efficiency. Those are Amazon’s stated goals; the article does not provide independent performance results.

Graviton4 is aimed at inference

Amazon’s other announcement, Graviton4, is an Arm-based processor intended for inference: running models after they have been trained. It is the fourth generation in Amazon’s Graviton family and is separate from the company’s Inferentia chip, which is also used for inference.

Compared with Graviton3, Amazon claims Graviton4 provides up to 30% better compute performance, 50% more cores, and 75% more memory bandwidth. The comparison is specifically with Graviton3, not the newer Graviton3E.

Amazon also says the physical hardware interfaces on Graviton4 are “encrypted.” The article says the precise meaning of that claim was unclear, and Amazon had been asked to explain it. The stated security detail therefore needs more context before customers can assess what it means for their workloads and data.

Graviton4 availability starts with preview

Graviton4 is planned for Amazon EC2 R8g instances. Those instances were available in preview at the time of the announcement, with general availability planned in the coming months.

Together, the two chips address separate computing needs: Trainium2 is for building models, and Graviton4 is for running them. Amazon’s claims describe the intended capacity and improvements, while availability schedules and further detail on the hardware encryption would help customers judge how the processors fit their plans.