paper-with-me

홈 › Papers

Adaptive Loss Balancing for Noise-Robust GRPO in Generative Recommendation

2026-06-07 · Kewei Xu, Junbo Qi, Yanyan Zou, Pengfei Zhang, Xingzhi Yao, Shengjie Li arxiv

Reinforcement learning (RL) presents a promising avenue for enhancing generative recommendation beyond supervised imitation, leveraging reward signals to guide policy improvement. However, its efficacy is critically contingent on the trustworthiness of the reward model for the samples it evaluates. In practice, production rankers, the widely adopted reward models, are trained on exposure-biased logs, leading to sample-dependent inaccuracies that violate this assumption. Our stratified analysis uncovers a consistent pattern: reward guidance is most beneficial when the policy exhibits uncertainty and the ranker can effectively discriminate the ground-truth item from rollout negatives. On other samples, the reward signal is either negligible or detrimental, highlighting the risk of uniform RL application. To address such an issue, we introduce AdaGRPO, a novel framework that treats reward-guided optimization as selective admission rather than uniform pressure. Training is anchored in supervised negative log-likelihood, while the GRPO objective is gated by a binary, per-sample clip determined by two rollout diagnostics: policy-side difficulty and reward discriminability. Instances failing either diagnostic default to pure supervision, ensuring stability and mitigating the amplification of noisy gradients. We validate AdaGRPO on a large-scale e-commerce dataset. At the best intermediate checkpoint, it elevates HR@10 from 11.01% to 12.18% while constraining hallucination below 0.22%, and maintains robustness at the final checkpoint (HR@10 11.63%, hallucination 0.27%), outperforming fixed NLL--GRPO mixtures across the retrieval--validity frontier. In production A/B tests, AdaGRPO achieves statistically significant gains in click-through rate and dwell time, confirming its practical utility.

📄 PDF Abstract BibTeX arXiv:2606.08480

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models

2025-10-20 · Jiajun Fan, Tong Wei, Chaoran Cheng, Yuxin Chen 외 arxiv

Balancing exploration and exploitation during reinforcement learning fine-tuning of generative models presents a critical challenge, as existing approaches rely on fixed divergence regularization that creates an inherent…

Text-to-Image GenerationReinforcement LearningStyle Transfer

Multi-Loss Rebalancing Algorithm for Monocular Depth Estimation

2020-08-01 · ECCV 2020 8 · Jae-Han Lee, Chang-Su Kim

An algorithm to combine multiple loss terms adaptively for training a monocular depth estimator is proposed in this work. We construct a loss function space containing tens of losses. Using more losses can improve infere…

Depth EstimationMonocular Depth Estimation

GAPO: Robust Advantage Estimation for Real-World Code LLMs

2025-10-22 · Jianqing Zhang, Zhezheng Hao, Wei Xia, Hande Dong 외 arxiv

Reinforcement learning (RL) is widely used for post-training large language models (LLMs) in code editing, where group-relative methods, such as GRPO, are popular due to their critic-free and normalized advantage estimat…

Reinforcement Learning

Hierarchy-Consistent Learning and Adaptive Loss Balancing for Hierarchical Multi-Label Classification

2025-08-19 · Ruobing Jiang, Mengzhe Liu, Haobing Liu, Yanwei Yu arxiv

Hierarchical Multi-Label Classification (HMC) faces critical challenges in maintaining structural consistency and balancing loss weighting in Multi-Task Learning (MTL). In order to address these issues, we propose a clas…

Hierarchical Multi-label ClassificationContrastive LearningMulti-Task Learning

TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models

2025-12-09 · Zheng Ding, Weirui Ye arxiv

Reinforcement learning (RL) post-training is crucial for aligning generative models with human preferences, but its prohibitive computational cost remains a major barrier to widespread adoption. We introduce \textbf{Tree…

Reinforcement Learning