paper-with-me

홈 › Papers

Riemann-Bench: A Benchmark for Moonshot Mathematics

2026-04-08 · Suhaas Garre, Erik Knutsen, Sushant Mehta, Edwin Chen arxiv

Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competition-style problem solving. However, competition mathematics represents only a narrow slice of mathematical reasoning: problems are drawn from limited domains, require minimal advanced machinery, and can often reward insightful tricks over deep theoretical knowledge. We introduce Riemann-Bench, a private benchmark of expert-curated problems designed to evaluate AI systems on research-level mathematics that goes far beyond the olympiad frontier. Problems are authored by Ivy League mathematics professors, graduate students, and PhD-holding IMO medalists, and routinely took their authors weeks to solve independently. Each problem undergoes double-blind verification by two independent domain experts who must solve the problem from scratch, and yields a unique, closed-form solution assessed by programmatic verifiers. We evaluate frontier models as unconstrained research agents, with full access to coding tools, search, and open-ended reasoning, using an unbiased statistical estimator computed over 100 independent runs per problem. Our results reveal that all frontier models currently score below 10%, exposing a substantial gap between olympiad-level problem solving and genuine research-level mathematical reasoning. By keeping the benchmark fully private, we ensure that measured performance reflects authentic mathematical capability rather than memorization of training data.

📄 PDF Abstract BibTeX arXiv:2604.06802

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics

2025-05-06 · Junqi Liu, Xiaohan Lin, Jonas Bayer, Yael Dillies 외

Neurosymbolic approaches integrating large language models with formal reasoning have recently achieved human-level performance on mathematics competition problems in algebra, geometry and number theory. In comparison, c…

Benchmarking

MOONSHOT : A Framework for Multi-Objective Pruning of Vision and Large Language Models

2026-04-14 · Gabriel Afriat, Xiang Meng, Shibal Ibrahim, Hussein Hazimeh 외 arxiv

Weight pruning is a common technique for compressing large neural networks. We focus on the challenging post-training one-shot setting, where a pre-trained model is compressed without any retraining. Existing one-shot pr…

Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study

2024-08-26 · Liuchang Xu, Shuo Zhao, Qingming Lin, Luyao Chen 외

The emergence of large language models such as ChatGPT, Gemini, and others highlights the importance of evaluating their diverse capabilities, ranging from natural language understanding to code generation. However, thei…

8kBenchmarkingCode GenerationNatural Language Understanding

A Quest for Knowledge

2021-02-26 · Christoph Carnehl, Johannes Schneider

Is more novel research always desirable? We develop a model in which knowledge shapes society's policies and guides the search for discoveries. Researchers select a question and how intensely to study it. The novelty of …

Decision Making

The geometry of the deep linear network

2024-11-13 · Govind Menon

This article provides an expository account of training dynamics in the Deep Linear Network (DLN) from the perspective of the geometric theory of dynamical systems. Rigorous results by several authors are unified into a …