How SparseGPT could shrink large language models

SparseGPT uses one-shot pruning to remove less important model parameters while preserving nearly all accuracy in the reported results. The researchers say it can make models such as OPT-175B 50 to 60 percent sparse, with pruning taking about four hours on a single GPU.

WTF Index NEUTRAL
◄ Terminator 0 Idiocracy 0 ►

The story describes a model compression technique that reduces resource demands while preserving accuracy, without a clear Terminator or Idiocracy lean.

How SparseGPT could shrink large language models

Large language models can demand substantial memory and computing power. SparseGPT offers a way to reduce that burden by removing parameters that contribute less to a model’s output, while aiming to keep its accuracy intact.

Why model size is a practical barrier

The article describes GPT-175B as a model with 175 billion parameters requiring at least 320 gigabytes of memory. Operating it therefore takes a minimum of five A100 GPUs, each with 80 gigabytes of memory.

One common compression route is quantization, which uses a less precise numerical representation for individual weights. That can make a network smaller, but the reduced precision can also affect performance.

Pruning takes a different route. It removes redundant or less important information from a model, making the remaining network more compact. The challenge has been maintaining accuracy: pruning can reduce it, and recovering the loss may require costly retraining.

A one-shot approach to pruning

SparseGPT is designed to prune a model in one pass, rather than relying on a lengthy cycle of pruning and retraining. The method is presented by Elias Frantar and Dan Alistarh from the Institute of Science and Technology Austria in a paper titled “Massive Language Models Can Be Accurately Pruned in One-Shot”.

The authors describe SparseGPT as the first precise one-shot pruning method that works efficiently on models with ten to 100 billion parameters. The article reports that the method can also be applied to the largest publicly available GPT models mentioned: OPT-175B and BLOOM-176B.

For those models, the team said pruning took about four hours with a single GPU. That runtime matters because earlier one-shot pruning methods were too time-consuming for models with billions of parameters. A technique that can handle them efficiently could make pruning a more practical compression option.

Removing parameters while preserving accuracy

The researchers report that SparseGPT reduced models by 50 to 60 percent. In the case of OPT-175B, they found virtually no accuracy loss compared with the dense model, even at that level of sparsity.

In practical terms, the article says about 100 billion parameters in OPT-175B could be ignored during inference. This illustrates the central idea behind sparse modeling: a model may contain many parameters that can be removed without a meaningful change in its measured accuracy.

The results suggest a way to reduce the computational demands of using very large language models. A smaller effective model could require fewer resources to run, though the source does not specify a new hardware requirement or quantify any resulting savings.

What the researchers want to explore next

The team believes progressive pruning combined with fine-tuning could reach at least 80 to 90 percent sparsification. That is a projection, rather than a result established in the reported experiments.

The researchers also plan to study whether their approach can be used during training. If it can reduce the computational cost of pre-training massive models, sparsification could affect the resources needed to build models as well as those needed to run them.

The article places this work alongside a sparsification approach demonstrated by German AI startup Aleph Alpha and British AI chip manufacturer Graphcore in November 2022. Together, these efforts point to interest in leaner language models. SparseGPT’s reported results offer one possible route: remove parameters selectively, then assess whether the model retains its accuracy.