paper-with-me

홈 › Papers

RL in the Wild: Characterizing RLVR Training in LLM Deployment

2025-09-29 · Jiecheng Zhou, Qinghao Hu, Yuyang Jin, Zerui Wang, Peng Sun, Yuzhe Gu, Wenwei Zhang, Mingshu Zhai, Xingcheng Zhang, Weiming Zhang arxiv

Large Language Models (LLMs) are now widely used across many domains. With their rapid development, Reinforcement Learning with Verifiable Rewards (RLVR) has surged in recent months to enhance their reasoning and understanding abilities. However, its complex data flows and diverse tasks pose substantial challenges to RL training systems, and there is limited understanding of RLVR from a system perspective. To thoroughly understand the system challenges introduced by RLVR, we present a characterization study of RLVR tasks in our LLM deployment. Specifically, we investigate the distribution and variation trends of workloads across different RL tasks across training steps. We identify issues such as GPU idling caused by skewed sequence length distribution, inefficient parallel strategies in dynamically varying workloads, inefficient data management mechanisms, and load imbalance. We describe our observations and call for further investigation into the remaining open challenges. Furthermore, we propose PolyTrace benchmark suite to conduct evaluation with realistic workloads, and a practical use case validates that PolyTrace benchmark suite exhibits 94.7% accuracy.

📄 PDF Abstract BibTeX arXiv:2509.25279

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs

2026-05-18 · Hamed Khosravi, Xiaoming Huo arxiv

A local specialist LLM, fine-tuned with reinforcement learning from verifiable rewards (RLVR) on operator-local data, is installed in a regulated organization with per-deployment error budget $α$. The operator needs a sa…

Reinforcement Learning

CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback

2026-06-01 · Bin Chen, Xinye Liao, Yiming Liu, Xin Liao 외 arxiv

Recent LLM search agents use reinforcement learning with verifiable rewards (RLVR) to learn search-augmented reasoning from outcome rewards. On hard problems, these agents rarely sample end-to-end successful rollouts, le…

Reinforcement Learning

Improving Generalization Robustness of Multimodal RLVR

2026-08-09 · Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng, Chenrui Zhou 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, whic…

Reinforcement Learning

Aletheia: What Makes RLVR For Code Verifiers Tick?

2026-01-17 · Vatsal Venkatkrishna, Indraneil Paul, Iryna Gurevych arxiv

Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training. However, their adoption in code generation has lagged behind that of execution …

Reinforcement LearningCode Generation

Language Models that Think, Chat Better

2025-09-24 · Adithya Bhaskar, Xi Ye, Danqi Chen arxiv

Reinforcement learning with verifiable rewards (RLVR) improves language model reasoning by using rule-based rewards in verifiable domains such as mathematics and code. However, RLVR leads to limited generalization for op…

Reinforcement LearningGeneral Knowledge