Fine-tuning pushes Code Llama ahead on HumanEval

A fine-tuned 34B Code Llama model reported a 67.6 percent HumanEval score, slightly above the 67 percent score GPT-4 achieved when released in March. The result highlights how open-source teams can adapt a released model, though benchmark scores alone do not describe every programming task or real-world use.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story reports a coding benchmark improvement from fine-tuning, with no clear lean toward AI control or human dependence.

Fine-tuning pushes Code Llama ahead on HumanEval

A fine-tuned version of Meta’s Code Llama reached a higher HumanEval score than the GPT-4 result reported at its March release. The comparison shows how quickly outside teams can adapt an open-source coding model, while also raising a practical question: what does a benchmark lead tell us about everyday coding work?

Fine-tuning changes the comparison

Phind, an AI co-programming startup, reported results for fine-tuned 34B variants of Code Llama. In the first run, its standard model scored 67.6 percent on HumanEval, while its Python model scored 69.5 percent.

The source compares those results with GPT-4’s 67 percent score on the same benchmark when it was released in March. The original 34 billion-parameter Code Llama scored 48.8 percent, according to Meta, and its Python variant scored 53.7 percent.

Those figures describe different versions and stages of development. The Phind models had been adapted for programming tasks, so the comparison is not simply between untouched releases. It instead illustrates how fine-tuning can shift a model’s performance on a defined evaluation.

Training on programming examples

Phind says it fine-tuned the two models natively on a custom dataset of about 80,000 high-quality programming tasks and solutions. That task-focused material is central to the reported improvement: the models were refined using examples directly related to the kind of work HumanEval measures.

The article also says Meta had already fine-tuned Code Llama to a 62 percent HumanEval success rate, according to Phind. Meta used 15,000 examples to refine Unnatural Code Llama, while Phind used about 80,000 examples for its models. These counts offer context for the approaches described, but they do not by themselves establish which training choices caused the score differences.

Phind reported training the models with 32 A100-80 GB GPUs and a sequence length of 4096 tokens, over three hours. The researchers used DeepSpeed ZeRO 3 and Flash Attention 2 for faster and more efficient training.

Later results add perspective

An update dated Aug 26, 2023 described another fine-tuned model from the WizardLM developers: WizardCoder-34B. It scored 73.2 percent in the first pass of HumanEval, according to the article.

WizardLM also noted that the GPT-4 version tested through the API at the end of August achieved a HumanEval score of 82, while GPT-3.5 scored 72.5. These later figures put the earlier comparison in perspective: a model’s benchmark result depends on which version and test run are being compared.

HumanEval is presented as an important evaluation for AI programming tasks. A score can help compare performance on that test, but the results here do not establish how a model performs across every coding problem or in practical use. The benchmark is one part of the picture.

Open access brings opportunity and limits

Phind published both models on Huggingface under the Llama license. The source says the license allows scientific and commercial use, with a special license required for use in widespread applications. It also says data generated with Llama 2 may not be used to train new AI models.

That licensing context matters alongside the scores. Open availability gives developers and researchers a way to work with the models, while the stated conditions shape how they may be used. The article describes a broader cycle in which community refinements can improve released models in benchmarks, potentially helping Meta improve its models faster through open-source development.

For readers evaluating coding models, the key takeaway is to look beyond a headline score. The model’s version, fine-tuning, benchmark run, and license all affect what the result means and how the model can be applied.