How Phi-1 Made Better Code by Learning from Better Data

Microsoft’s Phi-1 showed that carefully selected training data can help a small model perform strongly on coding benchmarks. The Python-focused model still has limits, including weaker handling of unfamiliar APIs, varied coding styles, and errors in prompts.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story describes a coding model improved through curated training data, with no clear lean toward harm or human dependence.

How Phi-1 Made Better Code by Learning from Better Data

A language model’s training does not depend only on its size or the amount of material it reads. Microsoft’s Phi-1 research points to a third factor: the quality of that material. The small model performed strongly on coding benchmarks after training on a curated mix of code and generated examples.

Building a training set with care

The researchers trained Phi-1 using what they described as “textbook quality” data. They filtered code from The Stack and StackOverflow datasets, selecting six billion high-quality training tokens with a classifier based on GPT-4. They added another billion tokens generated using GPT 3.5.

The model’s training took about four days on eight Nvidia A100 graphics cards. The researchers’ approach focused on preparing a useful collection of examples rather than simply increasing the volume of training data.

That distinction matters because a large dataset can contain material that is repetitive or less useful for the task at hand. The Phi-1 work suggests that selecting examples for clarity and quality can help a model learn coding patterns efficiently. It does not show that data quantity is irrelevant; instead, it highlights quality as another important part of training.

Strong benchmark results, with a narrow focus

The largest small model in the work, phi-1 1.3B, was further refined with code tasks. It outperformed models ten times larger that used 100 times more data on the HumanEval and MBPP benchmarks. In those test scenarios, only GPT-4 scored ahead of phi-1.

The results surpassed the researchers’ expectations. They linked the performance to the training data and titled their paper “Textbooks is all you need,” echoing the title of Google’s Transformer research, “Attention is all you need.” The comparison captures the paper’s central idea: well-chosen instructional examples may help a model build useful skills.

Those benchmark results describe a particular kind of performance. Phi-1 is specialized in Python programming, so its results do not establish that it can match larger models across every subject or coding task.

What the model still struggles with

Phi-1’s specialization limits its versatility. The researchers note that it lacks some of the domain-specific knowledge found in larger language models, including knowledge needed to program with specific APIs.

Its structured training also leaves it less robust when prompts vary in style or contain input errors. This is a practical boundary on what benchmark success means: performance on coding tests does not guarantee that the model will handle every real-world instruction equally well.

The synthetic examples also had imperfections. The team said GPT 3.5, used to generate some of the data, had a high error rate. They suggested that generating synthetic data with GPT-4 could allow further performance improvements. Even with errors in the generated material, however, Phi-1 learned effectively and produced correct code. The researchers said this indicates that a model can extract useful patterns from faulty data.

Why data quality is hard to measure

The researchers argue that high-quality training data is critical, but assembling it is challenging. A useful collection needs balance and diversity, and should avoid repetition. The article notes that measurement methods are lacking, especially for diversity and repetition.

That leaves a broader challenge for AI training: teams need ways to judge not just how much data they have, but whether the examples cover a suitable range of material without unnecessary duplication. Phi-1 offers a case study in how a carefully prepared dataset can support a specialized model, while its limitations show that quality alone does not remove the need to consider coverage and reliability.

Phi-1 was expected to be released as open source on Hugging Face. The work also fits a view shared by Andrei Karpathy, former head of AI at Tesla and now back at OpenAI, who expects more “small and powerful expert models” trained with attention to data quality, diversity, and complementary synthetic data.