A smaller language model does not always need to imitate a larger model’s answers to learn from it. Microsoft researchers say Orca 2 can benefit by learning how to approach a task, choosing among reasoning strategies such as breaking a problem into steps or giving a direct answer.
The approach is intended to improve what compact models can do while using less computing power than large models. The results are promising, the researchers report, though Orca 2 still has limitations including hallucinations and content errors.
Teaching a model how to approach a task
Many efforts to build smaller models rely on imitation learning: a smaller system reproduces outputs from a larger one. The Microsoft team argues that copying a teacher’s answers or style may limit a smaller model’s potential, because the approach that works for a powerful model may not suit a less capable one.
Orca 2 instead learns from the teacher model’s reasoning process. The researchers call this “explanation tuning.” They created an extended, customized synthetic dataset that teaches a range of methods for generating responses.
Those methods include working through a problem step by step, recalling information before generating an answer, combining recall with reasoning, extracting relevant information before generating, and responding directly. The idea is to let the model use a strategy suited to the task, rather than apply one approach everywhere.
Why a smaller model may need a different strategy
A capable model such as GPT-4 may be able to answer a complex question directly. A smaller model may do better if it separates the work into steps. Orca 2’s training aims to teach that distinction: the most useful path to an answer can depend on the model and the task.
The training examples came from a more powerful teacher model. For this experiment, Microsoft used GPT-4 via ChatGPT. The researchers say the teacher’s quality matters to how well the method works, and describe the results as potentially state-of-the-art and an upper limit for what Orca can currently achieve.
Orca 2 is based on Meta’s LLaMA 2 model family. The work fits into a broader effort to develop generative AI models that can deliver useful performance with lower computing demands. Microsoft has also recently demonstrated that interest with Phi-2.
How the researchers evaluated Orca 2
The team assessed the model on 15 benchmarks covering approximately 100 tasks and more than 36,000 individual test cases. The evaluations used zero-shot scenarios, meaning the model was tested without examples of the specific task in the prompt.
The benchmarks covered language comprehension, everyday knowledge, multi-level thinking, mathematical problem-solving, reading comprehension, summarizing, grounding, truthfulness, and identifying toxic content. Together, they were designed to examine a broad range of capabilities rather than a single kind of response.
According to the researchers, Orca 2 substantially outperformed models of similar size. On some tasks, its performance was comparable to or better than models five to ten times larger, particularly on complex tasks that test advanced reasoning in zero-shot settings.
Promising results, with familiar limitations
The findings do not mean that Orca 2 removes the weaknesses associated with language models. The researchers note distortions, limited transparency, hallucinations, and content errors. They also say the model may inherit many limitations from its teacher.
They see potential to improve reasoning, control, and safety by using synthetic data during post-training. The broader goal is not simply to replace large foundational models, which the team says will continue to demonstrate superior capabilities. Smaller systems could instead support applications where efficiency and performance need to be balanced for a particular deployment.
Microsoft is making Orca 2 available on Hugging Face as open source for research purposes. The work offers one example of how training a model to select a solution strategy could help smaller systems handle more demanding tasks.