Mathematical Reasoning
10개 벤치마크 · 논문 1,929편 · 이 태스크의 논문 보기 →
Benchmarks
AIME24
FrontierMath
Lila (IID)
Lila (OOD)
PGPS9K
AMC23
GeoQA
MATH500
UniGeo
UniGeo (PRV)
Most implemented
LoRA: Low-Rank Adaptation of Large Language Models
Analysing Mathematical Reasoning Abilities of Neural Models
Motif 3: Technical Report
Qwen2.5 Technical Report
Mistral 7B
Training Verifiers to Solve Math Word Problems
Papers
From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are oft…
Mathematical ReasoningWhich Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen al…
Mathematical ReasoningDifficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the mode…
Mathematical ReasoningReinforcement LearningRISE: Recursive Improvement via Self-Extrapolating Policy Distillation
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while se…
Mathematical ReasoningCode GenerationDo LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory
Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of its prerequisites, a principle formalized by Knowledge Space Theory (KST). While LLMs achieve strong performance…
Mathematical ReasoningFirst Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves
Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potential to act as autonomous agents for real-world tasks. However, scenarios requiring agents to fulfill users' com…
Reinforcement LearningMathematical Reasoning