Running an AI model can consume energy each time it processes an input, even after the model has been trained. IBM’s NorthPole processor targets that repeated work: it is designed to execute certain neural networks while reducing the energy spent moving data between memory and computing hardware.
A processor built for inference
NorthPole focuses on inference, the stage when a trained neural network handles information such as an image or an audio clip. It does not reduce the energy required to train a neural network, and it is not a general-purpose AI processor.
That focus sets limits on the tasks it can handle. The article notes that large language models may be too large to fit in NorthPole’s hardware. The chip’s efficiency comes from matching its design to inference workloads, rather than supporting every kind of AI computation.
NorthPole also differs from neuromorphic hardware. It draws on ideas from IBM’s earlier TrueNorth work, but its processing units perform calculations; they do not imitate the spiking signals used by actual neurons.
Keeping data close to the calculations
In conventional processors and GPUs, neural network weights are stored in memory and must be brought to the units that perform calculations. Moving that information consumes energy. NorthPole addresses this by placing local memory and computing capacity together in each unit, so weights can stay near where they are used.
The chip arranges its computational units in a 16×16 array. Its on-chip networks move results between units, distribute the weights and code for each layer, and support communication between neighboring units. That local communication can help with image tasks where nearby pixels need to be considered together, such as identifying an object’s edge.
Simpler calculations, more parallel work
Each unit is optimized for lower-precision calculations, from two- to eight-bit precision. Inference often does not need the same precision as training, so the chip can use simpler arithmetic for its intended workloads.
The units also cannot make conditional branches based on variable values. In practical terms, code for the chip cannot include an “if” statement. Removing that capability avoids hardware for speculative branch execution and helps keep the units focused on parallel calculations.
At two-bit precision, each unit can perform over 8,000 calculations in parallel. IBM’s team also developed training software to determine the minimum precision needed at each layer for the network to work successfully.
Promising results, with limits
NorthPole’s test chips were made using a 12 nm process. A chip with 22 billion transistors contains 256 computational units, each with 768 kilobytes of memory. In tests against an Nvidia V100 Tensor Core GPU made using a similar process, NorthPole performed 25 times the calculations for the same amount of power. The article also reports that it could outperform a cutting-edge GPU by about fivefold on that measure.
Those results came from a research prototype installed on a PCIe card. IBM told Ars that more work would be needed before turning it into a commercial product, and the company did not say whether it planned to do so.
There are practical constraints, too. A neural network must fit within the chip’s hardware; a layer with too many nodes can exceed its capacity. Splitting layers across multiple NorthPole chips and running those parts in parallel is a possible approach, but it had not been tested.
NorthPole demonstrates how specialized hardware can reduce power use for a defined category of AI work. Its results do not establish that one processor can efficiently handle all AI tasks: the design’s gains depend on inference workloads that match its architecture and fit within its resources.