Language models typically rely on tokenizers to turn input into units a model can process. Meta AI’s MegaByte explores a different route: working with bytes, the basic pieces of digital text and other data, while using a layered architecture to manage long sequences.
Why look beyond tokenization?
A tokenizer may represent a short word as one token and a longer word as several. That makes language model input manageable, but the approach has tradeoffs. Depending on the model architecture, token processing can be computationally intensive, adding new kinds of input can be difficult, and the model may not handle individual letters directly.
That last gap can show up in small tasks, such as counting the letter “n” in “mayonnaise.” A model that processes words as larger units may not naturally track each character. More broadly, tokenization adds a separate processing stage between the original input and the model.
Long inputs raise another challenge. Books, videos, and podcasts can produce very large sequences. The source reports that GPT-4 or Claude models could handle between 32,000 and 100,000 tokens, while byte-level processing can create even longer sequences that are difficult to manage efficiently.
How MegaByte handles bytes
MegaByte avoids a classical tokenizer and works with text, images, and audio at the byte level. It first divides a sequence into patches. A patch embedder then represents each patch by losslessly concatenating embeddings for its individual bytes, such as letters.
The work is split between two kinds of transformer. A global autoregressive module processes the patch representations and passes them along. A local autoregressive model handles each patch and predicts the bytes inside it.
This division lets the model work with detailed byte-level information while organizing the sequence into larger chunks. In principle, the global component can operate across patches while the local component focuses on the contents of each one.
What the early comparisons show
Meta says the architecture supports more computational parallelism, allows larger and more powerful models at the same computational cost, and reduces the cost of the transformers’ self-attention mechanism. Those are proposed efficiency benefits of the design, rather than evidence that every larger model will achieve them.
The researchers compared MegaByte with a simple decoder-transformer architecture and Deepmind’s PerceiverAR on tasks involving text, images, and audio. They reported that MegaByte was more efficient and could handle sequences approaching a million bytes. Meta’s description also says it outperformed existing byte-level models across a range of tasks and modalities, with competitive language modeling results against subword models.
These findings point to a possible way to reduce reliance on tokenization. They do not establish that MegaByte can already replace tokenizers in large, deployed language models: the experiments used models well below the size of current language models.
The scaling question remains
OpenAI’s Andrej Karpathy called the approach “promising” and wrote, “Everyone should hope we can throw away tokenization in LLMs.” He also cautioned that a straightforward byte-level approach produces sequences that are too long, leaving the implementation details as the central challenge.
Meta AI likewise framed the results as an indication that MegaByte may have the potential to replace classic tokenizers in transformer models. The team’s next step, according to the source, was to scale the experiments to much larger models and datasets.
That next step matters because strong results on smaller models do not settle how the method will behave at larger scales. MegaByte offers a structured way to process bytes and promising early efficiency results; its broader value depends on whether those gains hold as model size and training data grow.