paper-with-me

홈 › Papers

IsarStep: a Benchmark for High-level Mathematical Reasoning

2020-06-13 · ICLR 2021 1 · Wenda Li, Lei Yu, Yuhuai Wu, Lawrence C. Paulson

A well-defined benchmark is essential for measuring and accelerating research progress of machine learning models. In this paper, we present a benchmark for high-level mathematical reasoning and study the reasoning capabilities of neural sequence-to-sequence models. We build a non-synthetic dataset from the largest repository of proofs written by human experts in a theorem prover. The dataset has a broad coverage of undergraduate and research-level mathematical and computer science theorems. In our defined task, a model is required to fill in a missing intermediate proposition given surrounding proofs. This task provides a starting point for the long-term goal of having machines generate human-readable proofs automatically. Our experiments and analysis reveal that while the task is challenging, neural models can capture non-trivial mathematical reasoning. We further design a hierarchical transformer that outperforms the transformer baseline.

📄 PDF Abstract BibTeX arXiv:2006.09265

Code (2)

reactive-systems/circuit-repair tf
reactive-systems/ml2

Tasks

Mathematical ProofsMathematical ReasoningVocal Bursts Intensity Prediction

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Residual Connection 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Multi-Head Attention 설명 없음
Adam 설명 없음
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

2024-10-10 · Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai 외

Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e…

GSM8KMathMathematical Reasoning

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

2025-03-27 · Haoxiang Sun, Yingqian Min, Zhipeng Chen, Wayne Xin Zhao 외

In recent years, the rapid development of large reasoning models has resulted in the saturation of existing benchmarks for evaluating mathematical reasoning, highlighting the urgent need for more challenging and rigorous…

Data VisualizationMathMathematical Reasoning

PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts

2025-04-25 · Yiming Wang, Pei Zhang, Jialong Tang, Haoran Wei 외

In this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, an…

DiversityMathematical Reasoning

LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches

2026-04-02 · Linyang He, Qiyao Yu, Hanze Dong, Baohao Liao 외 arxiv

Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligence and cognitive science. As LLMs are in…

Mathematical Reasoning

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

2025-01-23 · Xin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao 외

Large Language Models (LLMs) have made significant strides in mathematical reasoning, underscoring the need for a comprehensive and fair evaluation of their capabilities. However, existing benchmarks often fall short, ei…

Mathematical Reasoning