UC Berkeley researchers used feedback from AI models to train Starling-7B, an open language model built on Openchat 3.5 and, before that, Mistral-7B. Their work explores whether AI-generated preferences can help make chatbot replies more helpful and safe.
The results suggest that this approach can improve how a model responds, though they do not show a broad gain in its underlying abilities. Starling-7B still struggles with reasoning and mathematics, can hallucinate, and is vulnerable to jailbreaks.
How AI feedback shaped the model
Reinforcement Learning from AI Feedback, or RLAIF, uses evaluations from AI systems to guide the training of another model. It differs from Reinforcement Learning from Human Feedback, or RLHF, where people rate model outputs. Human ratings were used to improve GPT-3.5 and GPT-4, helping make interactions with ChatGPT feel more natural.
For Starling-7B, the researchers created Nectar, a dataset containing 183,000 chat prompts and seven responses for each prompt. Those responses produced 3.8 million pairwise comparisons. They came from several models, including GPT-4, GPT-3.5-instruct, GPT-3.5-turbo, Mistral-7B-instruct, and Llama2-7B.
GPT-4 scored the synthetic responses. The researchers also designed a method to address GPT-4's tendency to give higher scores to the first and second responses it saw. This matters because the training signal depends on comparisons being useful, not simply on one response appearing earlier in a list.
Benchmark gains, with a narrow focus
The team evaluated Starling-7B with MT-Bench and AlpacaEval, benchmarks that use GPT-4 to score performance on simple instruction-following tasks. These evaluations focus on qualities such as helpfulness and safety, so their scores offer a view of response behavior rather than a complete measure of what a model can do.
On MT-Bench, Starling-7B scored 8.09, compared with 7.81 for vanilla Openchat 3.5. Its AlpacaEval result rose from 88.51% to 91.99%. The researchers reported that Starling-7B outperformed most models on MT-Bench, with OpenAI's GPT-4 and GPT-4 Turbo ahead. On AlpacaEval, its results were on par with commercial chatbots such as Claude 2 or GPT-3.5.
The researchers say RLAIF mainly improved helpfulness and safety. Knowledge questions, mathematics, and coding remained about the same or were minimally degraded. In other words, the reported gains center on the way the model handles instructions, rather than a general jump in its core capabilities.
What the results leave open
Benchmark scores can only say so much about how a chatbot will perform in practice. Here, GPT-4 judged both the training responses and performance on the cited evaluations. Human raters may have preferences that differ from GPT-4's, so a strong score under this setup does not establish that people would consistently prefer the model's answers.
The researchers describe the findings as evidence that RLAIF can work when GPT-4's preferences serve as the reward signal. They also point to adding high-quality human feedback to Nectar as a possible next step. That could help make the training preferences better reflect what people want from a chatbot.
Starling-7B retains familiar limitations of language models. It can produce hallucinations, has difficulty with reasoning or mathematics tasks, and is vulnerable to jailbreaks because it was not explicitly trained for those scenarios. Improvements in helpfulness and safety therefore should not be read as proof that these risks have been resolved.
Research release and availability
The researchers are publishing the Nectar dataset, the Starling-RM-7B-alpha reward model trained with it, and the Starling-LM-7B-alpha language model on Hugging Face under a research license. They said code and a paper would follow shortly. The model can also be tested in the chatbot arena.
The release gives researchers and other users a way to examine an approach that substitutes AI evaluations for human ratings during part of the training process. Its benchmark results are promising for RLAIF, while the difference between GPT-4 preferences and human preferences remains an important question for further work.