Meta’s OPT-IML adapts its OPT language model to follow instructions across a broad set of language tasks. The release includes two versions trained on different portions of a collection of about 2,000 tasks, along with benchmarks for evaluating them. Access is limited to research use: the provided license does not allow commercial use or redistribution.
Training OPT to follow task instructions
OPT-IML stands for “Open-Pre-trained-Transformer - Instruction Meta-Learning.” It builds on Meta’s OPT language model, announced in early May 2022 and released in late May. Meta describes the IML version as fine-tuned to perform better on natural language tasks than the original OPT.
The task mix includes familiar uses of language models, such as answering questions, summarizing text, and translating. Researchers drew on about 2,000 natural language tasks and organized them into eight NLP benchmarks, collectively called OPT-IML Bench. Meta also provides those benchmarks so performance can be assessed across tasks.
The release separates training and evaluation in one of its versions. OPT-IML was trained on 1,500 tasks, with another 500 held back for evaluation. OPT-IML-Max used all 2,000 available tasks for training. That distinction matters when interpreting results: the first version includes an explicit set of withheld tasks, while the second is trained on the full collection.
Reported improvements depend on the task
Meta reports that OPT-IML improves over native OPT by approximately 6-7% on average in 0-shot accuracy at both the 30B and 175B model scales. In 0-shot evaluation, the model is assessed without examples in the prompt. For 32-shot accuracy, Meta reports significant improvements on the 30B model and milder improvements on the 175B model.
The averages do not mean that instruction-tuning helps every benchmark equally. Meta says gains are significant on tasks including RTE, WSC, BoolQ, ARC, CB, and WiC. It also reports that performance does not improve on StoryCloze, PIQA, Winograd, and Winogrande.
This uneven picture is useful context for people assessing the release. OPT-IML is presented as a more capable model for language tasks overall, but its results depend on the benchmark and evaluation setup. A single average cannot show where the model performs better or where it does not.
Benchmarks test different kinds of generalization
The researchers’ paper describes evaluation splits intended to examine three kinds of generalization: performance when tasks are fully supervised, performance on unseen tasks from categories represented during training, and performance on tasks from categories kept entirely out of training.
These distinctions help separate learning a set of tasks from applying instruction-tuning beyond those tasks. A model may perform well on tasks it has seen while producing different results on new tasks or unfamiliar categories. The evaluation suite is designed to make those differences visible and, according to the paper, supports tradeoffs and recommended practices for instruction-tuning.
For readers comparing instruction-tuned language models, the benchmark design offers a reminder to look beyond headline accuracy. Whether the evaluation uses familiar tasks, new tasks in familiar categories, or entirely held-out categories changes what the results can demonstrate.
Research access comes with restrictions
Meta is making both 30-billion-parameter OPT-IML versions available as downloads on Github. The 175B version is planned to be available by request, with a request form also planned for Github. The largest model has 175 billion parameters, the same size as OpenAI’s GPT-3. The source says the original OPT was reported to have incurred only one-seventh the CO₂ footprint of GPT-3 during training.
The release is not for commercial deployment. The OPT license is limited to non-commercial research, is bound to the recipient, and cannot be redistributed. That means access to the model and its reported task improvements should be understood within those conditions.
OPT-IML adds instruction-tuning and a task-focused evaluation suite to Meta’s OPT family. Its benchmark results show gains on some language tasks, while also documenting cases where tuning does not improve performance. For researchers, both the task coverage and the license terms shape how the release can be used.