paper-with-me

홈 › Papers

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

2026-05-25 · Li Wang, Xiaodong Lu, Xiaohan Wang, Yikun Ban, Jiajun Chai, Wei Lin, Tianhao Peng, Guojun Yin arxiv

Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewards (RLVR). Nonetheless, RLVR intrinsically relies on ground-truth labels for reward computation, the acquisition of which is often prohibitively expensive in real-world scenarios. While unsupervised RLVR paradigms attempt to circumvent this by training on pseudo-labels, they are notoriously susceptible to training collapse. Moreover, different samples often exhibit varying annotation values. In this paper, we propose Reinforcement Learning with Active Verifiable Rewards (RLAVR), which actively acquires ground-truth labels for a small set of selected samples and integrates them with pseudo-labels, thereby stabilizing training dynamics and improving performance under limited annotation budgets. To identify valuable samples, we propose the Corrective Advantage Gap (CAG) metric and analyze the sample-level supervision value. Building on this, we introduce Correction-Aware Reliability Estimation for RLAVR (CARE), which translates the oracle CAG criterion into a practical pre-query acquisition policy to substantially improve training stability. Extensive experiments across diverse domains, model families, and model scales demonstrate the effectiveness and generality of our approach. Our code is available at https://github.com/Lumina04/CARE.

📄 PDF Abstract BibTeX arXiv:2605.25864

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?

2026-02-04 · Pingyue Zhang, Zihan Huang, Yue Wang, Jieyu Zhang 외 arxiv

Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capacity for active, self-directed exploratio…

Users Mispredict Their Own Preferences for AI Writing Assistance

2026-01-08 · Vivian Lai, Zana Buçinca, Nil-Jana Akpinar, Mo Houtti 외 arxiv

Proactive AI writing assistants need to predict when users want drafting help, yet we lack empirical understanding of what drives preferences. Through a factorial vignette study with 50 participants making 750 pairwise c…

Belief sharing: a blessing or a curse

2024-07-02 · Ozan Catal, Toon Van de Maele, Riddhi J. Pitliya, Mahault Albarracin 외

When collaborating with multiple parties, communicating relevant information is of utmost importance to efficiently completing the tasks at hand. Under active inference, communication can be cast as sharing beliefs betwe…

Self-Guided Belief Propagation -- A Homotopy Continuation Method

2018-12-04 · Christian Knoll, Adrian Weller, Franz Pernkopf

Belief propagation (BP) is a popular method for performing probabilistic inference on graphical models. In this work, we enhance BP and propose self-guided belief propagation (SBP) that incorporates the pairwise potentia…

Exploiting Context for Robustness to Label Noise in Active Learning

2020-10-18 · Sudipta Paul, Shivkumar Chandrasekaran, B. S. Manjunath, Amit K. Roy-Chowdhury

Several works in computer vision have demonstrated the effectiveness of active learning for adapting the recognition model when new unlabeled data becomes available. Most of these works consider that labels obtained from…

Active LearningDocument ClassificationGeneral ClassificationScene Classification