Meta’s Humpback explores a way for a large language model to get better at following instructions using text that has not been labeled for that purpose. The method, called instruction backtranslation, has the model create and assess its own training examples instead of relying on human annotations or distillation from more powerful models.
Turning unlabeled text into instruction examples
Instruction backtranslation has two parts: self-augmentation and self-curation. First, the model uses an unlabeled text corpus to generate candidate instruction-response pairs. Given a piece of text, it tries to work backward and predict what instruction could have prompted that response.
This produces a collection of examples that can be used for fine-tuning. The examples are synthesized by the model, so the process does not begin with people writing an instruction for every response.
The model filters its own training data
Generating examples is only the first step. In self-curation, the model scores the candidate pairs, ranks them, and keeps the highest-scoring subset. Lower-quality candidates are filtered out before they become part of the selected training data.
The generation and selection steps then repeat. After an iteration, the improved model can produce better candidate examples and make stronger judgments about which demonstrations are useful. In this way, the model contributes both to creating its training material and to choosing which material to learn from.
The approach is a form of iterative self-training: each round uses the model’s current abilities to improve the data for a later round. The intended result is progress in two related skills—producing better instructions and recognizing higher-quality demonstrations.
Humpback’s results on instruction following
Meta’s researchers report that the approach improves instruction-following performance compared with previous work using a LLaMa model of the same scale. Their Humpback 65B model achieved state-of-the-art results among non-distilled LLaMa methods on the Alpaca instruction-following benchmark.
The source says Humpback surpassed models including Claude, Guanaco, LIMA, and Falcon-Instruct on that comparison. The result is specific to the benchmark and comparison described: it shows how the method performed among non-distilled LLaMa approaches, rather than establishing a general ranking across every model or task.
Why the data loop matters
Instruction tuning depends on examples that connect a request to a useful response. Backtranslation offers a way to build those examples from a larger pool of text, while self-curation tries to concentrate training on the strongest candidates. Because the same model participates in both stages, its ability to generate and evaluate examples can improve from one iteration to the next.
Meta’s team says it plans to scale the method by considering larger unlabeled corpora, which its analysis suggests could bring further gains. That next step follows the method’s central premise: more unlabeled text could give the model additional material from which to construct and select instruction examples.