paper-with-me

홈 › Papers

Putnam-like dataset summary: LLMs as mathematical competition contestants

2025-09-29 · Bartosz Bieganowski, Daniel Strzelecki, Robert Skiba, Mateusz Topolewski arxiv

In this paper we summarize the results of the Putnam-like benchmark published by Google DeepMind. This dataset consists of 96 original problems in the spirit of the Putnam Competition and 576 solutions generated by LLMs. We analyze the performance of models on this set of problems to verify their ability to solve problems from mathematical contests. We find that top models, particularly Gemini 2.5 Pro, achieve high scores, demonstrating strong mathematical reasoning capabilities, although their performance was lower on problems from the 2024 Putnam competition. The analysis highlights distinct behavioral patterns among models, including bimodal scoring distributions and challenges in providing fully rigorous justifications.

📄 PDF Abstract BibTeX arXiv:2509.24827

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs

2025-08-05 · Aryan Gulati, Brando Miranda, Eric Chen, Emily Xia 외 arxiv

Current mathematical reasoning benchmarks for large language models (LLMs) are approaching saturation, with some achieving > 90% accuracy, and are increasingly compromised by training-set contamination. We introduce Putn…

Mathematical Reasoning

PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition

2024-07-15 · George Tsoukalas, Jasper Lee, John Jennings, Jimmy Xin 외

We present PutnamBench, a new multi-language benchmark for evaluating the ability of neural theorem-provers to solve competition mathematics problems. PutnamBench consists of 1692 hand-constructed formalizations of 640 t…

Automated Theorem Proving

An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems

2025-08-12 · Yuren Hao, Xiang Wan, ChengXiang Zhai arxiv

In this paper, we introduce a systematic framework beyond conventional method to assess LLMs' mathematical-reasoning robustness by stress-testing them on advanced math problems that are mathematically equivalent but with…

Mathematical Reasoning

Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving

2025-07-03 · Matthieu Zimmer, Xiaotong Ji, Rasul Tutunov, Anthony Bordg 외 arxiv

Reasoning remains a challenging task for large language models (LLMs), especially within the logically constrained environment of automated theorem proving (ATP), due to sparse rewards and the vast scale of proofs. These…

Automated Theorem Proving

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

2026-03-20 · Guillaume Baudart, Marc Lelarge, Tristan Stérin, Jules Viennot arxiv

We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant, autonomously proved 10 of 12 problems from the 2025 Putnam Mathematical Compe…