paper-with-me

홈 › Papers

Limits of Generalization in RLVR: Two Case Studies in Mathematical Reasoning

2025-10-30 · Md Tanvirul Alam, Nidhi Rastogi arxiv

Mathematical reasoning is a central challenge for large language models (LLMs), requiring not only correct answers but also faithful reasoning processes. Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising approach for enhancing such capabilities; however, its ability to foster genuine reasoning remains unclear. We investigate RLVR on two combinatorial problems with fully verifiable solutions: \emph{Activity Scheduling} and the \emph{Longest Increasing Subsequence}, using carefully curated datasets with unique optima. Across multiple reward designs, we find that RLVR improves evaluation metrics but often by reinforcing superficial heuristics rather than acquiring new reasoning strategies. These findings highlight the limits of RLVR generalization, emphasizing the importance of benchmarks that disentangle genuine mathematical reasoning from shortcut exploitation and provide faithful measures of progress. Code available at https://github.com/xashru/rlvr-seq-generalization.

📄 PDF Abstract BibTeX arXiv:2510.27044

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

2026-07-16 · Yuxuan Zhu, Rohan Alur, Daniel Kang arxiv

While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In…

Reinforcement Learning

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping

2025-10-29 · Tue Le, Linh Ngo Van, Trung Le arxiv

Reinforcement learning with verifiable rewards (RLVR) has become a practical route to improve large language model reasoning, and Group Relative Policy Optimization (GRPO) is a widely used optimizer in this setting. Howe…

Reinforcement LearningMathematical ReasoningQuestion Answering

DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning

2026-02-18 · Haoxiang Sun, Lizhen Xu, Bing Zhao, Wotao Yin 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has been shown effective in enhancing the visual reflection and reasoning capabilities of Large Multimodal Models (LMMs). However, existing datasets are predominantly…

Reinforcement LearningMultimodal Reasoning

Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning

2025-02-27 · Sheng Zhang, Qianchu Liu, Guanghui Qin, Tristan Naumann 외

Reinforcement learning from verifiable rewards (RLVR) has recently gained attention for its ability to elicit self-evolved reasoning capabilitie from base language models without explicit reasoning supervisions, as demon…

MathMedical Question AnsweringMultiple-choiceMultiple Choice Question Answering (MCQA)+2

Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards

2025-11-21 · Zhen Wang, Zhifeng Gao, Guolin Ke arxiv

Test-time scaling has been shown to substantially improve large language models' (LLMs) mathematical reasoning. However, for a large portion of mathematical corpora, especially theorem proving, RLVR's scalability is limi…

Reinforcement LearningMathematical Reasoning