paper-with-me

홈 › Papers

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR

2026-04-06 · Kazuki Egashira, Mark Vero, Jasper Dekoninck, Florian E. Dorner, Robin Staab, Martin Vechev arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs). While RLVR is designed for tasks with verifiable ground-truth answers, real-world verifiers (e.g., static code checkers) can introduce errors into the reward signal. Prior analyses have largely treated such errors as random and independent across samples, concluding that errors merely slow training with limited effect on final performance. However, practical verifiers tend to exhibit systematic errors. This introduces a risk of models learning unwanted consistent behavior from a structurally incorrect reward signal. In this work, we study the impact of such systematic verification errors on RLVR. Through controlled experiments on arithmetic tasks, we show that systematic false negatives lead to similar effects as random noise. On the other hand, systematic false positives can cause a wide range of behaviors from sub-optimal plateaus to performance collapse. Crucially, these outcomes are not determined by the overall error rate but by the specific pattern of introduced errors, making pre-hoc mitigation difficult. Our results show that, in contrast to prior conclusions, realistic verification errors can critically shape RLVR outcomes and that verifier quality has to be understood beyond its sample-level error rate.

📄 PDF Abstract BibTeX arXiv:2605.02909

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Marginals Before Conditionals

2026-03-10 · Mihir Sahasrabudhe arxiv

We construct a minimal task that isolates conditional learning in neural networks: a surjective map with K-fold ambiguity, resolved by a selector token z, so H(A | B) = log K while H(A | B, z) = 0. The model learns the m…

What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers

2025-06-16 · Pulkit Gopalani, Wei Hu

Training Transformers on algorithmic tasks frequently demonstrates an intriguing abrupt learning phenomenon: an extended performance plateau followed by a sudden, sharp improvement. This work investigates the underlying …

The Effects of Communication Delay on Human Performance and Neurocognitive Responses in Mobile Robot Teleoperation

2025-08-25 · Zhaokun Chen, Wenshuo Wang, Wenzhuo Liu, Yichen Liu 외 arxiv

Communication delays in mobile robot teleoperation adversely affect human-machine collaboration. Understanding delay effects on human operational performance and neurocognition is essential for resolving this issue. Howe…

Collapse Grammar Optimizer: GH-Seed Trace Suppression Architecture

2025-04-30 · FlameSovereign Trace Grammar Release 2025 4 · FlameSovereign

A grammar-based optimizer preventing collapse through GH-trace lifecycle gating. This submission presents FlameSovereign's GHv1.0 seed, demonstrating trace stability under entropy pulse, adversarial spikes, and loss plat…

The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior

2026-03-30 · Laura Gomezjurado Gonzalez arxiv

Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the source of that delay remains poorly understood. In encoder-decoder arithm…