Code Checks Lift GPT-4 Math Accuracy to 84.3%

Researchers found that GPT-4 Code Interpreter scored 69.7% on the MATH benchmark, above GPT-4’s 42.2% and the previous state-of-the-art 53.9%. Two methods that prompt the system to verify its work with code raised its accuracy to 84.3%.

WTF Index TERMINATOR
◄ Terminator 1 Idiocracy 0 ►

The story describes a more capable AI system, with no clear societal harm or loss of human skill.

Code Checks Lift GPT-4 Math Accuracy to 84.3%

GPT-4 Code Interpreter performed better on a challenging math benchmark when researchers encouraged it to use code more often and check its answers. Their experiments suggest that code execution and verification can help the system catch and correct errors in its solutions.

Code execution changes the results

The researchers evaluated GPT-4 Code Interpreter, referred to as GPT4-Code, on mathematical reasoning datasets, including MATH. They describe MATH as the most challenging mathematical problem set. On that benchmark, GPT4-Code achieved 69.7% accuracy, compared with 42.2% for GPT-4 and a previous state-of-the-art score of 53.9%.

The team varied how prompts constrained the frequency of code use. Their findings linked GPT4-Code’s stronger performance to several abilities working together: generating code, executing it, assessing the output, and revising a solution when the result seemed unreasonable. In this setup, code served both as a way to work through problems and as a means of checking answers.

Two ways to make verification count

Building on that finding, the researchers tested two methods designed to encourage more frequent code execution. The first, Explicit Code-Based Self-Verification, asks the system to check its answer with code. If verification fails, it continues trying until the check succeeds.

The second method, Verification-Guided Weighted Majority Voting, uses verification results when combining candidate answers. Answers that pass verification receive greater weight, so the final choice reflects whether a proposed answer was checked successfully.

With these methods, GPT4-Code reached 84.3% accuracy on MATH, up from 69.7%. The researchers say that pushing the system to use code more often, and making its verification ability part of the process, was central to the improvement.

Results extend beyond one benchmark

The team also evaluated the approach on MMLU, using its math and science problems. The methods improved GPT-4 Code Interpreter’s accuracy across all datasets in that evaluation, according to the source article. This suggests the techniques may apply beyond the MATH problem set, though the reported findings concern the datasets the researchers tested.

The results focus on a particular system and prompting approach. They show how a model’s access to code execution can be used to test its own work, but they do not establish that the same gains will appear in every model or task.

Possible uses for other language models

The researchers plan to apply their findings about code-use frequency and the two verification methods to language models beyond GPT-4. They also want to use the approach to create more accurate datasets, with detailed step-by-step code-based solution generation and code-based validation.

Such datasets could help improve open-source models like LLaMA 2, the researchers suggest. The potential value is in pairing worked solutions with a check that uses code, giving future training data both a reasoning path and a way to validate the result.