paper-with-me

홈 › Papers

Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning

2026-04-23 · Qinan Yu, Alexa Tartaglini, Peter Hase, Carlos Guestrin, Christopher Potts arxiv

Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that the reasoning chains trained through RLVR reliably represent how a model gets to its answer. In this paper, we develop two metrics for critically examining this assumption: Causal Importance of Reasoning (CIR), which measures the cumulative effect of reasoning tokens on the final answer, and Sufficiency of Reasoning (SR), which measures whether a verifier can arrive at an unambiguous answer based on the reasoning alone. Through experiments with the Qwen2.5 model series and ReasoningGym tasks, we find that: (1) while RLVR does improve task accuracy, it does not reliably improve CIR or SR, calling the role of reasoning in model performance into question; (2) a small amount of SFT before RLVR can be a remedy for low CIR and SR; and (3) CIR and SR can be improved even without SFT by applying auxiliary CIR/SR rewards on top of the outcome-based reward. This joint reward matches the accuracy of RLVR while also leading to causally important and sufficient reasoning. These results show that RLVR does not always lead models to rely on reasoning in the way that is commonly thought, but this issue can be remedied with simple modifications to the post-training procedure.

📄 PDF Abstract BibTeX arXiv:2604.22074

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Future-as-Label: Scalable Supervision from Real-World Outcomes

2026-01-09 · Benjamin Turtel, Paul Wilczewski, Danny Franklin, Kris Skothiem arxiv

Time creates free supervision: forecasts about real-world events resolve to verifiable outcomes. The passage of time provides labels that require no annotation. To exploit this structure, we extend reinforcement learning…

Reinforcement Learning

Linear Combinatorial Semi-Bandit with Causally Related Rewards

2022-12-25 · Behzad Nourani-Koliji, Saeed Ghoorchian, Setareh Maghsudi

In a sequential decision-making problem, having a structural dependency amongst the reward distributions associated with the arms makes it challenging to identify a subset of alternatives that guarantees the optimal coll…

Decision MakingSequential Decision Making

Verifiable Process Rewards for Agentic Reasoning

2026-05-11 · Huining Yuan, Zelai Xu, Huaijie Wang, Xiangmin Yi 외 arxiv

Reinforcement learning from verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but most existing approaches rely on sparse outcome-level feedback. This sparsity creates a cred…

Reinforcement LearningLogical Reasoning

BeamPERL: Parameter-Efficient RL with Verifiable Rewards Specializes Compact LLMs for Structured Beam Mechanics Reasoning

2026-03-04 · Tarjei Paule Hage, Markus J. Buehler arxiv

Can reinforcement learning with hard, verifiable rewards teach a compact language model to reason about physics, or does it primarily learn to pattern-match toward correct answers? We study this question by training a 1.…

Reinforcement Learning

Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks

2025-09-29 · Peiran Xu, Zhuohao Li, Xiaoying Xing, Guannan Zhang 외 arxiv

Large Language Models (LLMs) increasingly rely on external tools such as search engines to solve complex agentic tasks that require reasoning and external knowledge retrieval. Recently, reinforcement learning with verifi…

Reinforcement Learning