LeanDojo gives language models a path into theorem proving

LeanDojo is an open-source platform that helps researchers connect language models with Lean, a proof assistant. Its tools support data extraction and model interaction, while the ReProver system retrieves mathematical premises to guide proof strategies.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

LeanDojo is a research tool for model-assisted theorem proving, with no clear lean toward harm or human dependence.

LeanDojo gives language models a path into theorem proving

Language models may be able to help prove mathematical theorems, but producing reliable proofs requires more than generating plausible text. LeanDojo is an open-source platform designed to connect these models with Lean, a proof assistant, giving researchers tools to study how automated proof work can be done.

Why theorem proving needs a different approach

Automated theorem proving, or ATP, aims to generate proofs for statements written in formal logic. It has applications in formal mathematics and can support formal verification, a way to check the correctness and security of high-risk applications.

One difficulty is the size of the search space: a system may need to explore many possible steps before finding a proof. Interactive theorem proving offers another route. In this approach, mathematicians work with software called proof assistants to construct proofs.

Machine learning could help automate parts of that interaction. Large language models paired with proof assistants such as Lean are one possible way to do so, but earlier methods have been difficult to reproduce or extend. A team of researchers from Caltech, Nvidia, MIT, UC Santa Barbara, and UT Austin identified proprietary code and data, along with high computational requirements, as barriers.

What LeanDojo provides

LeanDojo is intended to make learning-based theorem proving easier to investigate. It offers two central capabilities: extracting data and allowing models to interact programmatically with Lean.

The researchers describe it as the first tool that can reliably interact with Lean. That matters because the proof assistant can check whether a proposed proof step works, helping reduce errors in the interaction between a model and the formal system.

These capabilities give researchers a practical foundation for exploring model-assisted proofs. Data extraction can support the preparation of material for learning, while programmatic interaction gives a model a way to work with Lean rather than produce a proof in isolation.

ReProver uses retrieved premises

LeanDojo also addresses premise selection, a bottleneck in theorem proving. A proof strategy may depend on choosing useful statements from an existing mathematical library. ReProver, short for Retrieval-Augmented Prover, demonstrates an approach in which a language model generates a strategy using a small set of premises retrieved from Lean's math library.

The team reports that ReProver outperforms some other methods. It proves a significant percentage of theorems in the LeanDojo benchmark and in two existing datasets, MiniF2F and ProofNet. The source does not specify the percentages, so the results are best understood as evidence of progress rather than a claim that the system solves theorem proving broadly.

ReProver can also prove theorems that do not yet have a proof in Lean. That capability could help extend existing Lean math libraries by adding formal proofs for further results.

A benchmark for further research

The team is releasing a LeanDojo benchmark containing nearly 97,000 theorems from Mathlib, along with a plugin for ChatGPT. A shared benchmark gives researchers a common set of problems for evaluating systems and comparing future work.

The project arrives amid expectations that language models equipped with external tools could take on a larger role in mathematics. Mathematician Terence Tao recently predicted that such systems could become trusted co-authors in math and other sciences by 2026. LeanDojo and ReProver offer one example of the tools behind that possibility, while leaving room for others to test and improve the approach.