A growing set of open AI projects is using answers generated by ChatGPT as training material. Stanford’s Alpaca helped demonstrate the approach: a relatively small collection of AI-written examples could change how an existing language model responds to instructions.
Projects that followed adapted the idea for different models and uses. Their progress points to a practical route for experimentation, while licensing restrictions and the need for high-quality human data remain part of the picture.
Alpaca made the recipe visible
In mid-March, Stanford researchers introduced Alpaca, a version of Meta’s LLaMA 7B fine-tuned with 52,000 example statements generated by OpenAI’s GPT-3.5 (text-davinci-003). The team reported partially comparable results in its benchmarks.
The project drew attention in part because its reported cost was $600. Alignment researcher Eliezer Yudkowsky saw the result as a real challenge for companies like OpenAI. The broader lesson was that teams could use outputs from a capable AI system to help train another model to follow instructions.
Stanford shared the training data and code for generating the data and fine-tuning the model. It had not released Alpaca itself: LLaMA was not released for commercial use, and OpenAI’s GPT-3.5 terms prohibited using the model to develop AI models that compete with OpenAI.
Follow-on projects changed the ingredients
Other projects quickly built on or took inspiration from Alpaca. Alpaca-LoRA used the resource-efficient low-rank adaptation (LoRA) method with Meta’s LLaMA, aiming for results comparable to Alpaca. The method is also widely used in Stable Diffusion.
Nomic AI’s GPT4All took a different route through its data selection. It used 430,000 GPT-3.5-turbo outputs chosen from a dataset of one million outputs in total. That example illustrates how the training material itself can be a major part of a project’s design.
ChatDoctor focused on medical conversations. Its authors first trained LLaMA with the 52,000 Alpaca examples, then used 5,000 real conversations between medical professionals and patients. The sequence combined AI-generated instruction examples with human conversations tailored to a narrower purpose.
Databricks also applied the Alpaca training dataset, but paired it with EleutherAI’s GPT-J-6B rather than LLaMA. The company said that fine-tuning an older open model on a small corpus of instruction data could produce striking behavior. Together, these projects suggest that the approach was not tied to a single base model.
Reusable data offers reach, with limits
Alpaca’s influence came from making a process easier to reproduce: start with an existing model, prepare instruction examples using ChatGPT, and fine-tune the model. For open-source developers, this could make experimentation more accessible and support attempts to build free, efficient alternatives to ChatGPT.
But an accessible recipe does not automatically grant permission to use every component commercially. The LLaMA restriction and GPT-3.5 terms shaped what Stanford could release and how the work could be used. The source also notes that larger LLaMA models had been leaked, raising expectations for more capable models while leaving the commercial licensing issue unresolved.
AI-generated training data attracted attention beyond these open projects. A report said Google employees wanted to use ChatGPT dialogues to train Bard, but the effort stopped after an employee raised it with management. The episode underscored the perceived value of these conversations as a source of training examples.
Human data still matters
AI-generated examples may help teams build and refine models, but the source does not treat them as a full replacement for human work. High-quality human-generated data remains relevant for high-performance systems. OpenAI reportedly employs human experts to verify or create specialized material, including data for code tasks.
That leaves two complementary roles for training data. Generated examples can give open-source projects a usable starting point, while human expertise can help produce or check specialized material. The projects following Alpaca show how quickly the first role can spread; the continued use of human experts points to the work that remains important as models improve.