Microsoft’s LongNet Aims to Stretch AI Context to a Billion Tokens

Microsoft’s LongNet uses a modified attention method designed to handle sequences as long as a billion tokens while scaling compute linearly. The approach is a feasibility study so far, and its practical benefits still need to be demonstrated.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

LongNet is a feasibility study for expanding AI context, with no demonstrated societal harm or skill erosion.

Microsoft’s LongNet Aims to Stretch AI Context to a Billion Tokens

AI models can use longer context windows to process more text at once, but extending the amount of information they handle has a steep computational cost in standard Transformer architectures. Microsoft’s LongNet proposes a different way to manage that cost, with a design the team says could scale to a billion tokens.

Why longer context is difficult

A model’s sequence length sets how much information it can consider in one context. For a language model, that can mean processing and generating more text. For a vision Transformer, a longer sequence can let it capture more information from an image.

The obstacle is that, in a standard Transformer, compute requirements rise quadratically with sequence length. As the input grows, the resources needed to process it can quickly become a barrier. Researchers have explored ways to extend context, but LongNet takes aim at making the scaling relationship linear.

The article describes context windows ranging from 4,096 tokens for ChatGPT to about 100,000 tokens for a commercially available Claude model. Microsoft’s team says LongNet can scale to one billion tokens, or 250,000 times the ChatGPT context length cited in the article. That would correspond to about 750,000,000 words or 2,000,000 pages.

A different pattern of attention

LongNet’s central technique is an adapted attention mechanism called “dilated attention.” Attention helps a model relate different tokens in a sequence. LongNet changes how much attention it allocates according to the distance between those tokens.

Nearby tokens receive attention at a level comparable to standard attention. As tokens get farther apart, the attention allocation decreases exponentially and the model uses coarser patterns. The design aims to preserve close-up detail while reducing the effort spent connecting distant parts of very long sequences.

This tradeoff is the core of LongNet’s scaling proposal. The method does not treat every relationship across a vast input with the same level of detail. Instead, it makes local relationships more precise and distant relationships less fine-grained, with the goal of keeping compute growth manageable as sequence length increases.

What the team has demonstrated

In one test, the team trained a speech generation model with sequences of up to 32,000 tokens and compared it with classical Transformer-based approaches. The team reported that LongNet showed familiar scaling behavior: as the model grew, its perplexity decreased.

That result suggests the method can operate in a setting where model size and performance follow patterns already seen in classical Transformers. It does not, by itself, establish that billion-token context improves the quality of a language model or that the approach performs better on practical tasks.

The researchers say much longer context could make it possible to process datasets on the scale of the web. A model could have a larger memory and receptive field, and training data could contain longer and more complex causal or reasoning paths for it to learn from. The team sees these capabilities as possible routes toward models that generalize better.

Promises still to be tested

LongNet also offers a way to investigate the limits of in-context learning. The team says an extremely long context could help models with many-shot learning and may help alleviate catastrophic forgetting. These are potential benefits, rather than demonstrated outcomes in the reported test.

The paper, as described in the source article, does not compare LongNet with modern language models such as GPT-4 32k. It also lacks measures such as accuracy or human evaluations that could show whether the approach brings meaningful advantages to users. For now, LongNet is best understood as a feasibility study of very long sequences.

Whether a billion-token context will translate into better results remains an open question for follow-up work. The team plans to explore other applications, including multimodal large language models and genomic data modeling. Those directions could test whether the same attention design is useful beyond the initial speech generation experiment.