Technology companies are building larger computing systems to train generative AI models. Cerebras, Inflection AI and major cloud providers are adding capacity, giving developers and model builders access to hardware designed for demanding AI workloads.
Cerebras scales its Condor Galaxy system
Cerebras has introduced Condor Galaxy 1, a system rated at 2 Exaflops. It was assembled and put into operation in 10 days, using 32 Cerebras CS-2 computers. The system is owned by G42, an Abu Dhabi-based holding company whose businesses include G42 Cloud, one of the largest cloud computing providers in the Middle East.
Cerebras says Condor Galaxy is expected to double in size within the next 12 weeks. Over the next 18 months, the company plans to install additional systems, aiming for 36 Exaflops across 9 installations. That would make its network one of the largest supercomputing efforts focused on AI workloads.
The CS-2 systems use the Waferscale Engine-2, a processor designed for AI. Each chip is built from a single silicon wafer and contains 2.6 trillion transistors and 850,000 AI cores. The architecture is one approach to assembling substantial computing capacity for neural network training.
Inflection AI commits to a large GPU cluster
Inflection AI is building a supercomputer based on 22,000 Nvidia H100 GPUs. At a time when AI chip supply is limited, the company’s connections as an Nvidia investment target likely helped it secure the allocation, according to the source article. Inflection recently raised $1.3 billion from Microsoft, Nvidia and others.
The company has also introduced Inflection-1, its first proprietary language model. It is described as being on par with GPT 3.5, Chinchilla and PaLM-540B, and it powers Pi, Inflection’s personal AI. Inflection’s focus is on generating and understanding natural language, with the aim of letting people give computers complex tasks through conversation instead of learning to code.
CEO Mustafa Suleyman, a DeepMind co-founder, expects a breakthrough in conversational AI in the next five years. That expectation points to a product goal behind the hardware investment: building systems that can handle more useful interactions in everyday language.
Cloud services widen access to H100 GPUs
Large GPU systems are also becoming available through cloud platforms. AWS launched P5 instances that can run up to 20,000 H100 GPUs. Similar hardware is available through Microsoft Azure, Google and Core Weave.
Cloud access can make it quicker for developers to prototype generative AI applications, because they can obtain scaled computing resources without building an entire supercomputer themselves. The H100’s Transformer engine can also reduce training times, making large workloads more practical to run.
One benchmark example illustrates the speed possible with these resources: Core Weave and Nvidia trained a GPT-3 model with 175 billion parameters on about 3,500 GPUs in less than 11 minutes in the MLPerf benchmark. The result shows how hardware, software and access to many processors can combine to accelerate a model training run.
Capacity is becoming part of the AI race
The new systems reflect a shift in the scale of model development. Cerebras CEO Andrew Feldman said that the number of companies training neural network models with 50 billion or more parameters went from 2 in 2021 to more than 100 this year. That growth helps explain why companies are investing in clusters designed specifically for AI.
More computing capacity does not by itself determine what models will be built or how useful they will be. But it can make large training jobs easier to undertake, while cloud instances let more developers experiment with generative AI. Together, dedicated supercomputers and rented GPU capacity are expanding the infrastructure available for the next wave of models.