Google’s New TPUs Aim to Cut the Cost of AI Workloads

Google Cloud announced its fifth-generation tensor processing units for AI training and inferencing at Cloud Next. The company says the chips improve performance per dollar and let large workloads span multiple TPU clusters.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

This routine chip launch focuses on efficiency and scale without a clear lean toward harm or human dependence.

Google’s New TPUs Aim to Cut the Cost of AI Workloads

Google Cloud has introduced the fifth generation of its tensor processing units, or TPUs, for AI training and inferencing. Announced at Cloud Next, the company’s annual user conference, the new chips put efficiency and the ability to scale across clusters at the center of Google’s latest cloud computing offer.

Performance per dollar is the focus

Google says the fifth-generation TPUs deliver a 2x improvement in training performance per dollar compared with the previous generation. For inferencing, the company claims a 2.5x improvement in performance per dollar.

Those comparisons frame the announcement around the amount of work customers can get for their spending. The figures are Google’s stated improvements over the last generation; the source does not provide further details about how the comparisons were measured.

Mark Lohmeyer, the VP and GM for compute and ML infrastructure at Google Cloud, described the chip as “the most cost-efficient and accessible cloud TPU to date.” That statement reflects the company’s positioning for this release, while the performance-per-dollar comparisons offer its specific claims about efficiency.

Large AI workloads can span clusters

Google also says it has made it possible for customers to scale TPU clusters beyond what was previously possible. In Lohmeyer’s description, an AI workload can extend beyond the physical boundaries of one TPU pod or one TPU cluster and use multiple physical clusters.

He said a single large AI workload could scale to tens of thousands of chips. The practical point in the announcement is that customers are not limited to fitting a workload within a single physical cluster. Google presents this expanded reach as a way to support larger AI models and workloads.

The company links that scale with cost efficiency and customer choice. Lohmeyer said Google is offering options across cloud GPUs and cloud TPUs to meet the needs of a broad set of emerging AI workloads. The announcement does not specify which workloads are best suited to each processor type.

A broader set of Google Cloud compute options

The TPU launch came alongside a separate update about Nvidia hardware. Google said it would make Nvidia’s H100 GPUs generally available to developers the following month through its A3 series of virtual machines.

Together, the announcements point to Google Cloud offering both its custom TPUs and cloud GPUs for AI work. The source does not say that one replaces the other. Instead, Lohmeyer’s remarks emphasize flexibility and optionality across both kinds of cloud processors.

For developers, the news includes two distinct considerations: Google’s claimed efficiency gains for its new TPUs and the ability to grow workloads across multiple TPU clusters, alongside access to H100 GPUs through A3 virtual machines. The announcement supplies those broad capabilities, but leaves the choice between them to the needs of each workload.

From announcement to availability

Google had announced the fourth version of its custom processors in 2021, and those processors became available to developers in 2022. That earlier timeline offers context for the new generation, though the article does not give a specific availability date for the fifth-generation TPUs.

The H100 update does include a timing reference: Google said general availability for developers would begin next month. For the fifth-generation TPUs, the central details in the announcement are the stated performance-per-dollar improvements and the expanded ability to span clusters. Those claims describe Google’s intended value for customers considering cloud AI compute.