Generating text with a large language model takes substantial computation: the model predicts one word at a time, and each prediction must finish before the next begins. Google’s Confident Adaptive Language Modeling, or CALM, aims to make that process more efficient by spending less effort on words the model already predicts confidently.
Give each word the compute it needs
Language models use multiple transformer layers to process text. Attention and feedforward modules update the model’s internal representations, and the decoder uses those representations to predict the next word. In a standard setup, each word goes through the same full sequence of layers.
CALM changes that pattern. It checks how confident the model is in a prediction partway through the layers. If confidence is high enough, the model can use the current representation and move on. If confidence is low, it continues through later layers before settling on a prediction.
The approach reflects a practical distinction: some next-word predictions are straightforward, while others need more processing. CALM dynamically allocates computation across generation steps rather than applying the same amount to every word.
Why generation is costly
Large language models are used for tasks such as translation and text generation. Models including GPT-3, PaLM and LaMDA have demonstrated strong results, while the source notes that language model performance increases with model size.
That scale comes with a cost. Large models are slower than smaller variants and require substantial computation. And because generation proceeds one word at a time, the next prediction cannot begin until the current one is complete. CALM targets the work happening inside each prediction, allowing some words to exit the layer sequence early.
Tests report faster generation with maintained quality
Google tested the method by training a T5 model and comparing CALM with a standard model. The evaluation covered benchmarks for translation, summarization and question answering. Google reports that CALM achieved high benchmark scores while using fewer layers per word on average.
On TPUs, the method saved up to 50 percent of computation time in practice while maintaining quality. The reported result is a reduction in compute time, not a claim that every model, task or deployment will see the same savings. The source describes the outcome in the context of Google’s tests.
Maintaining output quality matters because reducing processing could otherwise affect predictions. CALM’s confidence check is intended to reserve deeper processing for cases where the model is less certain, while allowing confident predictions to use fewer layers.
A possible part of a broader efficiency toolkit
Google frames CALM as one way to make growing language models more efficient to use. The method adjusts the amount of computation dynamically during generation, so efficiency depends on whether a prediction can be made confidently before reaching the final layers.
Google also says CALM can be combined with other efficiency approaches, including distillation or sparsity. Together, these ideas point to different ways of reducing the resources needed to run large models. CALM’s particular contribution is to vary the computation from word to word, based on the model’s confidence.
For users, the potential benefit is faster text generation without a reduction in output quality in the reported tests. For model developers, the results suggest that not every generated word needs to consume the same processing budget. The tests offer evidence for that approach, while leaving its performance in other settings to be established.