Abstract
Theorem proving presents a significant challenge for large language models (LLMs) due to the requirement for formal proofs to be rigorously checked by proof assistants, such as Lean, eliminating any margin for error or hallucination. While existing LLM-based theorem provers attempt to operate autonomously, they often struggle with novel and complex theorems where human insights are essential. Lean Copilot is a novel framework that integrates LLM inference into the Lean proof assistant environment. In this work, we benchmark performance of several LLMs including general and math-specific models for theorem proving using the Lean Copilot framework. Our initial investigation suggests that a general-purpose large model like LLaMa-70B still has edge over math-specific smaller models for the task under consideration. We provide useful insights into the performance of different LLMs we chose for the task.
| Original language | English |
|---|---|
| Title of host publication | NLP4Science 2024 - 1st Workshop on NLP for Science, Proceedings of the Workshop |
| Editors | Lotem Peled-Cohen, Nitay Calderon, Shir Lissak, Roi Reichart |
| Publisher | Association for Computational Linguistics (ACL) |
| Pages | 208-218 |
| Number of pages | 11 |
| ISBN (Electronic) | 9798891761858 |
| DOIs | |
| State | Published - 2024 |
| Event | 1st Workshop on NLP for Science, NLP4Science 2024 - Miami, United States Duration: Nov 16 2024 → … |
Publication series
| Name | NLP4Science 2024 - 1st Workshop on NLP for Science, Proceedings of the Workshop |
|---|
Conference
| Conference | 1st Workshop on NLP for Science, NLP4Science 2024 |
|---|---|
| Country/Territory | United States |
| City | Miami |
| Period | 11/16/24 → … |
Funding
This research used resources of the Oak Ridge Leadership Computing Facility (OLCF), which is a DOE Office of Science User Facility at the Oak Ridge National Laboratory supported by the U.S. Department of Energy under Contract No. DEAC05-00OR22725. We would like to thank Dr. Tirthankar Ghosal for his guidance and support throughout the research. We would also like to thank Peiyang Song, for answering questions about experiments using Lean Copilot (Song et al., 2024). The code repository is available at: https://code.ornl.gov/v28/atp_lean_copilot.
Fingerprint
Dive into the research topics of 'Benchmarking Automated Theorem Proving with Large Language Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver