Meta Opens Llama 2 to More Builders, With Caveats

Meta made Llama 2 available for research and commercial use, with versions suited to chat and different model sizes. The company reports improved performance and helpfulness, while also acknowledging gaps in its benchmarks, biases and the uncertainty of how an openly available model may be used.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

Wider access to a capable model modestly raises misuse concerns, though the article mostly describes a routine release and acknowledges its limitations.

Meta Opens Llama 2 to More Builders, With Caveats

Meta has released Llama 2, a family of text-generating models that developers can use for research and commercial projects. The release makes the models available for fine-tuning through AWS, Azure and Hugging Face, widening access compared with the request-only availability of the original Llama.

A broader release, with several model options

The first Llama models could generate text and code from prompts, but access was limited. Meta said it gated them because of concerns about misuse; the model later leaked online and spread through AI communities.

Llama 2 is offered in two forms. The standard Llama 2 is a base model, while Llama 2-Chat has been fine-tuned for two-way conversation. Both come in 7 billion, 13 billion and 70 billion parameter versions. Parameters are learned parts of a model that help determine how it performs at tasks such as generating text.

Meta says the models are free for research and commercial use. It also points to easier deployment: the models are optimized for Windows through an expanded partnership with Microsoft, and the company says they are intended to run on smartphones and PCs using Qualcomm’s Snapdragon system-on-chip. Qualcomm said it was working to bring Llama 2 to Snapdragon devices in 2024.

More training, but limited visibility into the data

Meta says Llama 2 was trained on two trillion tokens, compared with 1.4 trillion for the earlier Llama. Tokens are pieces of text used in model training; a word can be split into several pieces. The article notes that larger training sets generally help generative AI, though the amount of data alone does not establish how a model will perform.

The company’s whitepaper describes the training material as web content, mostly in English, with an emphasis on text considered factual. Meta says it did not use data from its own products or services, but it does not identify the specific sources.

That lack of detail matters because the origin and use of training material are contested. The source article reports that thousands of authors signed a letter urging technology companies not to use their writing to train AI models without permission or compensation. Without more information about the dataset, outsiders have limited ability to assess what material shaped Llama 2.

Performance claims come with qualifications

Meta’s comparisons found Llama 2 somewhat behind prominent closed models GPT-4 and PaLM 2 on a range of benchmarks. The gap was more pronounced for computer programming. Meta also said human evaluators found Llama 2 roughly as “helpful” as ChatGPT across roughly 4,000 prompts designed to assess helpfulness and safety.

Those results are not a guarantee of performance in everyday use. Meta acknowledges that its tests cannot cover every real-world situation and may lack diversity, including enough coverage of coding and human reasoning. Benchmark results offer one view of a model, but the company itself identifies limits in what its evaluations measure.

The models also reflect biases in their training data. Meta says Llama 2 is more likely to generate “he” than “she” pronouns, and that its data has a Western skew, including an abundance of the words “Christian,” “Catholic” and “Jewish.” Toxic material in the training data also means it does not outperform other models on toxicity benchmarks.

Safety remains a deployment question

Llama 2-Chat performs better than the standard Llama 2 on Meta’s internal helpfulness and toxicity benchmarks, according to the company. But the chat models can be overly cautious: they may decline certain requests or provide more safety detail than needed.

Hosted versions may add protections beyond the models themselves. Meta said its Microsoft collaboration uses Azure AI Content Safety to detect “inappropriate” content in AI-generated images and text, with the aim of reducing toxic outputs from Llama 2 on Azure.

Meta’s whitepaper also says users must follow its license, acceptable use policy and guidance on safe development and deployment. Still, making models widely available means their eventual uses are hard to predict. Llama 2 gives more developers access to a capable text-generation system, while leaving questions about dataset transparency, model limitations and real-world use for developers and the broader public to weigh.