paper-with-me

홈 › Papers

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning

2026-06-11 · Akshay Krishnamurthy, Audrey Huang, Nived Rajaraman arxiv

Reinforcement learning has rapidly emerged as a key component in the training of reasoning and coding models, yet it remains poorly understood from a mechanistic perspective. We study how and through what underlying processes capabilities are acquired or enhanced via reinforcement learning post-training. Our analysis, based on controlled math reasoning experiments with Qwen-2.5-1.5B, reveals two core mechanisms: strategy selection and strategy improvement. Our results highlight the role of SFT data and reinforcement learning data in activating these mechanisms, in particular showing how supervising the model on diverse reasoning strategies can enable strategy selection and how increasing difficulty in reinforcement learning data can enable strategy improvement. Taken together, our results provide mechanistic insight into RL training and suggest practical interventions to continue scaling reasoning capabilities.

📄 PDF Abstract BibTeX arXiv:2606.13125

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking

2026-02-24 · Muhammad Saif Ullah Khan, Didier Stricker arxiv

Modeling spinal motion is fundamental to understanding human biomechanics, yet remains underexplored in computer vision due to the spine's complex multi-joint kinematics and the lack of large-scale 3D annotations. We pre…

What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models

2025-07-09 · Keyon Vafa, Peter G. Chang, Ashesh Rambachan, Sendhil Mullainathan

Foundation models are premised on the idea that sequence prediction can uncover deeper domain understanding, much like how Kepler's predictions of planetary motion later led to the discovery of Newtonian mechanics. Howev…

Inductive Bias

Multiphysics continuum mechanics models to advance glaucoma research: State-of-the-art and future perspectives

2024-11-10 · Daniel Sebastia-Saez, Jinyuan Luo, Mengqi Qin, Tao Chen 외

This review provides a broad vision on the use of mechanistic mathematical models based on continuum mechanics to tackle current challenges in glaucoma research. At present, the advent of Artificial Intelligence and data…

From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching

2026-03-27 · Yuyang Ji, Yixuan Shen, Shengjie Zhu, Yu Kong 외 arxiv

We present BioCoach, a biomechanics-grounded vision--language framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, through a novel three-stage pipeline: an exe…

Assessing Large Language Models in Mechanical Engineering Education: A Study on Mechanics-Focused Conceptual Understanding

2024-01-13 · Jie Tian, Jixin Hou, Zihao Wu, Peng Shu 외

This study is a pioneering endeavor to investigate the capabilities of Large Language Models (LLMs) in addressing conceptual questions within the domain of mechanical engineering with a focus on mechanics. Our examinatio…

Multiple-choicePrompt Engineering