paper-with-me

홈 › Papers

IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation

2025-09-30 · Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck, Jeremy Feusi, Tim Gehrunger, Raphael Appenzeller, Pieter Belmans, Alessio Bottini, Jim Bryan, João Camarneiro, Ana Cannas da Silva, Niklas Canova, Ana-Maria Castravet, Timo de Wolff, Claudio Fontanari, Filippo Gaia, Baran Hashemi, Daniel Holmes, David Holmes, Aitor Iribar Lopez, Victor Jaeck, Martina Jørgensen, Steven Kelk, Martijn Kool, Stefan Kuhlmann, Adam Kurpisz, Johannes Lengler, Chiara Meroni, Ingmar Metzler, Martin Möller, Samuel Muñoz-Echániz, David Muñoz-Lahoz, Robert Nowak, Georg Oberdieck, Daniel Platt, Dylan Possamaï, Gabriel Ribeiro, Aluna Rizzoli, Daria Sakhanda, Raúl Sánchez Galán, Zheming Sun, Diaaeldin Taha, Josef Teichmann, Richard P. Thomas, Henk van der Pol, Michel van Garrel, Charles Vial, Ignacio Barros, Benjamin Doerr, Peter Grünwald, Henry Liu, David Martins, Aleksandar Mijatović, Sergej Monavari, Marc Roth, Patrick Schnider, Yannik Schuler, Pim Spelier, Yuuji Tanaka, Ronald van Luijk arxiv

As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of mathematical knowledge. However, existing benchmarks are limited, as they focus solely on final-answer questions or high-school competition problems. To address this gap, we introduce IMProofBench, a private benchmark consisting of 77 peer-reviewed problems developed by expert mathematicians. Each problem requires a detailed proof and is paired with subproblems that have final answers, supporting both an evaluation by human experts and a large-scale quantitative analysis through automated grading. Furthermore, unlike prior benchmarks, the evaluation setup simulates a realistic research environment: models operate in an agentic framework with tools like web search for literature review and mathematical software such as SageMath. Our results show that current LLMs can already solve a significant percentage of research-level questions. IMProofBench will continue to evolve as a dynamic benchmark in collaboration with the mathematical community, ensuring its relevance for evaluating the next generation of LLMs.

📄 PDF Abstract BibTeX arXiv:2509.26076

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

2026-05-09 · Zicheng Lyu, Wenjie Yang, Shengzhong Zhang, Zengfeng Huang arxiv

Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the literature. When a proof reaches a non-trivial step, a useful assist…

Mathematical Reasoning

RMA: an Agentic System for Research-Level Mathematical Problems

2026-05-20 · Zelin Zhao, Bo Yuan, Jaemoo Choi, Yongxin Chen arxiv

We present $\textbf{Research Math Agents (RMA)}$, an agentic framework for automated reasoning on research-level mathematical problems. Unlike prior studies centered on competition mathematics or formal theorem proving, …

Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory

2026-07-07 · Jihao Liu, Guoxiong Gao, Zeming Sun, Bin Wu 외 arxiv

Recent LLM-based mathematical reasoning agents have begun to tackle research-level problems and, in several cases, have contributed to the resolution of open problems. However, scaling and orchestrating such agents effec…

Mathematical Reasoning

Learning to Match Mathematical Statements with Proofs

2021-02-03 · Maximin Coavoux, Shay B. Cohen

We introduce a novel task consisting in assigning a proof to a given mathematical statement. The task is designed to improve the processing of research-level mathematical texts. Applying Natural Language Processing (NLP)…

ArticlesAutomated Theorem ProvingInformation RetrievalRetrieval

Mask-Proof: An LLM-based Automated Data Curation Pipeline on Mathematical Proofs

2026-06-13 · Jierui Zhang, Siyuan Tan, Xinhang Li, Longzhuangzhi Lin 외 arxiv

Large language models (LLMs) are increasingly capable of mathematical problem solving and can even assist with research-level proofs, yet we still lack a scalable and reproducible way to measure step-level reasoning in l…

Mathematical Reasoning