How FreeWilly 2 pushed Llama 2 further with less training data

Stability AI and CarperAI released FreeWilly models built on Meta’s Llama models. FreeWilly 2 uses Llama-2 and a synthetic instruction dataset; it scored about four points ahead of Llama v2 on average across benchmarks, though Llama 2 remained slightly ahead on MMLU.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 0 ►

This is a routine model fine-tuning update, with only a mild Terminator lean from improving AI capabilities.

How FreeWilly 2 pushed Llama 2 further with less training data

Stability AI and CarperAI introduced two open-access language models called FreeWilly. The newer model, FreeWilly 2, builds on Meta’s Llama-2 and offers an example of how fine-tuning and a purpose-built dataset can change a model’s performance.

Fine-tuning with a synthetic dataset

Both FreeWilly models are based on Meta’s Llama models. FreeWilly 2 uses the newer Llama-2 model, which has 70 billion parameters. The team describes its contribution as careful fine-tuning with a synthetic dataset made from high-quality instructions.

The approach draws on Microsoft’s Orca method. In that method, a smaller model learns step-by-step reasoning from training examples generated by a larger language model, rather than simply trying to imitate the larger model’s output style.

For the examples used in this process, Microsoft researchers used GPT-4 to create a dataset containing step-by-step reasoning processes. The broader aim is to transfer some capabilities from a larger model to a smaller one, following a teacher-student principle. Orca performed better than similarly sized models in some tests, although it did not match the original models.

A smaller training set

The FreeWilly team says it assembled 600,000 data points using its selected prompts and language models. That is about ten percent of the dataset used by the Orca team, according to the article.

A smaller dataset can mean less training is required. The team also says the reduced training effort improves the model’s environmental footprint. The report does not quantify that effect, so the claim is best understood as the team’s stated rationale for using a more compact dataset.

This setup helps explain what the FreeWilly work demonstrates: model development does not stop when a base model is released. Developers can apply additional training data and fine-tuning to adapt an existing model and then compare the results on established benchmarks.

Benchmark results show strengths and limits

On common benchmarks, the FreeWilly model trained with this approach achieved results on par with ChatGPT in some logical tasks. FreeWilly 2, which is based on Llama 2, clearly outperformed FreeWilly 1 in the reported comparisons.

Across all benchmarks, FreeWilly 2 was about four points ahead of Llama v2 on average. The result suggests that Meta’s standard model had room for improvement and that open-source development could contribute to that progress.

The overall lead has an important qualification. Llama 2 remained slightly ahead on MMLU, a general language understanding benchmark described as important in the source. The comparisons therefore point to a model with strong results across the reported set, while also showing that its performance varies by task.

Research access, with a non-commercial condition

FreeWilly 1 and FreeWilly 2 were released for research purposes under a non-commercial license. That condition shapes how the models can be used: the release supports research access, but the article does not describe it as permission for commercial use.

The release also illustrates how open-access model development can build on existing systems. FreeWilly 2 starts from Llama-2, uses a synthetic instruction dataset, and is assessed against other models through benchmark results. The gains reported by the team are promising, but the MMLU result is a reminder that an average score does not tell the whole story.