Nvidia's Eos supercomputer completed a benchmark GPT-3 training task in 3.9 minutes, using 10,752 H100 GPUs. Nvidia projects that a full run using a much larger training dataset could take eight days, putting the system's capacity in the range of training a model comparable in scale to GPT-3.5. That estimate is a projection, not a report that Eos trained GPT-3.5 itself.
What the benchmark measured
The results came from MLPerf training benchmark 3.1. In the reported test, Eos trained a GPT-3 model with 175 billion parameters on 1 billion tokens. Its hardware combined 10,752 H100 GPUs with Nvidia's Quantum-2 InfiniBand network.
The 3.9-minute result nearly tripled a prior record of 10.9 minutes. Nvidia had set that earlier mark less than six months before, using just under 3,500 H100 GPUs. The comparison suggests that the system could add computing power without losing much efficiency: tripling the GPU count produced a 2.8x performance increase, or 93 % efficiency. The source attributes some of the improvement to software optimizations.
Microsoft also submitted a result for a system with 10,752 H100 GPUs, using Azure HD H100 v5. It completed the same GPT-3 training benchmark in just under 4 minutes. These results offer a way to compare large systems under the benchmark's conditions, while leaving the full cost and practical details of an extended training run outside the reported figures.
Why the eight-day estimate is different
The benchmark used 1 billion tokens, while Nvidia's projection assumes training on 3.7 trillion tokens. That larger quantity comes from the optimal data amount described by Chinchilla's results for a modern GPT-3 model with 175 billion parameters. Nvidia says Eos could complete that projected run in eight days.
The distinction matters: the measured result is a short benchmark run, while the eight-day figure is an estimate based on scaling up the training data. It does not establish that Eos has trained a model with GPT-3.5's capabilities, or that it could reproduce the original model. The article says the amount of data OpenAI used for GPT-3.5 is unclear.
For context, the source reports that OpenAI trained GPT-3 with 300-500 billion tokens. It says GPT-4 is rumored to have used nearly 13 trillion tokens, and that the original GPT-3.5 likely falls somewhere between those amounts. It also notes that GPT-3.5-turbo appears to be a smaller model. Those comparisons show why the eight-day projection can suggest a GPT-3.5-scale training run without proving equivalence to any particular model.
Other results show that scaling varies
MLPerf 3.1 also included Stable Diffusion training for the first time. With 1,024 Nvidia H100 GPUs, the task took 2.5 minutes; with 64 H100 GPUs, it took 10 minutes. The article says diffusion-model training did not scale as efficiently as large language model training. Intel's Gaudi 2 took just under 20 minutes with 64 accelerators.
The benchmark results also showed Intel's Gaudi 2 making a significant performance improvement over its previous round, overtaking the A100 and coming closer to the H100 when training large language models. The article reported expectations that Gaudi 3 could arrive as early as 2024 and might catch up with Nvidia's accelerator in some areas. Together, these findings make clear that hardware comparisons depend on the task being measured.
How to read the results
MLPerf's purpose is to provide transparent, objective tests that help users make informed purchasing decisions. Its supporting organizations include Amazon, Arm, Baidu, Google, Harvard, HPE, Intel, Lenovo, Meta, Microsoft, Nvidia, Stanford University, and the University of Toronto.
For buyers, the Eos result is evidence of how Nvidia's GPUs, network, and software performed on a defined training benchmark. The projected eight-day run points to what that system might do at a larger scale. It should be read as a capacity estimate based on the stated assumptions, rather than a verified training time for GPT-3.5.