More Training Data Lets Meta’s LLaMA Compete With Larger Models

Meta introduced four LLaMA language models, ranging from 7 to 65 billion parameters, trained on at least 1T tokens. The researchers say some smaller versions can match or outperform much larger systems on many tasks, while requiring less computing power to run.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story reports a routine model-efficiency advance, with no clear lean toward harmful autonomy or human dependence.

More Training Data Lets Meta’s LLaMA Compete With Larger Models

Meta’s LLaMA models suggest that language model performance depends on more than parameter count. The research team introduced four models ranging from 7 to 65 billion parameters and said that training smaller models on more data can help them compete with much larger systems.

Smaller models, competitive results

Meta said its 13-billion-parameter LLaMA performs better than the open-source OPT model and GPT-3, which has 175 billion parameters, on “most” language tasks. The 65-billion-parameter version, meanwhile, is described by the researchers as competitive with Google’s 540-billion-parameter Palm and on par with Deepmind’s Chinchilla.

These comparisons come from the research team, and they do not mean that every model performs equally well at every task. They do show why parameter count alone offers an incomplete picture: how a model is trained also matters to its results.

Training with more data

Meta trained all four LLaMA models on at least 1T tokens, a larger amount of training data than is typical for models at this scale, according to the researchers. The 7-billion-parameter model was still improving after processing 1T tokens.

The approach drew inspiration from Deepmind’s Chinchilla, which also used more training data than usual. Meta’s researchers describe the comparison as significant because both efforts point to a way of improving performance by pairing models with larger training datasets.

That approach has a cost. More training takes time and computing resources, and the article says LLaMA required a similar number of training hours, and therefore a similar amount of CO₂, to 175-billion-parameter OPT and Bloom models. But the researchers argue that a smaller model trained for longer may be cheaper to use once it is serving requests.

Why running costs matter

Training and serving a language model are different stages, with different costs. Training produces the model; inference is the work of generating outputs when people use it. A model that takes longer to train can still be preferable if it is faster and less expensive to run repeatedly.

The article explains that scaling laws focused on balancing model size and dataset size within a training compute budget. That calculation does not account for inference costs, which become important when a model is used at scale. From that perspective, a smaller model trained on more data could be a practical choice even if training it takes substantial resources.

Meta’s researchers said the 13-billion-parameter LLaMA operates at GPT-3 level and can run on a single Nvidia Tesla V100 graphics card. They suggested that this could broaden access to large-scale language models and support more research.

Public data and access limits

Meta said LLaMA was trained only on publicly available data, unlike some other large language models that use undocumented or non-public datasets. A cleaned version of the English Common Crawl dataset supplies 67% of its data; other sources include public GitHub and Wikipedia.

The researchers described the data choices as “compatible with open-sourcing,” but the article raises questions about whether public availability alone amounts to consent for AI training. It also notes that common open-source licenses do not yet provide for such use, and that models typically do not cite sources in their outputs. The article says courts may help clarify the issue in the future.

Meta made the models available under a non-commercial GPL v3 license to selected partners in academia, government and industry. The team also said it planned to release larger models trained on larger pretraining corpora and to fine-tune models with instructions. Together, those plans point to continued work on both model scale and the amount and kind of data used in training.