paper-with-me

Mathematical Reasoning

10개 벤치마크 · 논문 1,929편 · 이 태스크의 논문 보기 →

Benchmarks

AIME24

결과 9개

FrontierMath

결과 6개

Lila (IID)

결과 6개

Lila (OOD)

결과 6개

PGPS9K

결과 6개

AMC23

결과 3개

GeoQA

결과 2개

MATH500

결과 1개

UniGeo

결과 1개

UniGeo (PRV)

결과 1개

Most implemented

Motif 3: Technical Report

2026-08-10 · 구현 6개

Qwen2.5 Technical Report

2024-12-19 · 구현 6개

Mistral 7B

2023-10-10 · 구현 6개

Papers

From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning

2026-09-09 · Weichen Dai, Rafael Medeiros Cabral, Ziyi Shou, Yan Cao 외 arxiv

Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are oft…

Mathematical Reasoning

Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning

2026-09-09 · Yaning Jia, Chunhui Zhang, Wenxuan Xu, Xingjian Diao 외 arxiv

Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen al…

Mathematical Reasoning

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

2026-09-08 · Youngjun Yu, Sanghwan Jang, Hwanjo Yu hf

Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the mode…

Mathematical ReasoningReinforcement Learning

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

2026-09-04 · Yang Li, Semih Yavuz, Shafiq Joty arxiv

On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while se…

Mathematical ReasoningCode Generation

Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory

2026-09-04 · Peng Cui, Heejin Do, Mrinmaya Sachan arxiv

Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of its prerequisites, a principle formalized by Knowledge Space Theory (KST). While LLMs achieve strong performance…

Mathematical Reasoning

First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves

2026-09-04 · Tianjie Ju, Xinyue Xu, Wanxuan Sun, Lingxiao Diao 외 arxiv

Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potential to act as autonomous agents for real-world tasks. However, scenarios requiring agents to fulfill users' com…

Reinforcement LearningMathematical Reasoning

전체 1,929편 보기 →