paper-with-me

홈 › Papers

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF

2026-06-25 · Arnav Raj arxiv

Reinforcement learning from human feedback (RLHF) in production does not always have a synchronous reward signal. Code-execution verifiers, slow judge ensembles, and queued human review can return several gradient steps after the rollout that produced them, breaking the synchronous-reward assumption underlying standard PPO. We address this gap with Retroactive Advantage Correction (RAC): each pending slow completion is queued, aged through a non-negative kernel, and reinjected as a clipped residual into the next optimiser step's advantage. We prove that under an unbiased clipped importance ratio, the cumulative RAC correction is exactly unbiased when the effective delay kernel reinjects all of its mass, and carries a bias linear in the unreinjected fraction otherwise; at the no-delay identity kernel it reduces to V-trace. On a tabular Markov decision process (MDP) proof-of-concept, RAC reduces the closed-form policy bias by up to 47.9x at the two-slow-channel configuration, beating wait-for-slow at lower wall-clock cost. RAC integrates with PPO and GRPO through a two-line reward-manager patch.

📄 PDF Abstract BibTeX arXiv:2606.27580

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving

2026-03-16 · Yihong Guo, Dongqiangzi Ye, Sijia Chen, Anqi Liu 외 arxiv

Autonomous driving requires safe planning, but most learning-based planners lack explicit self-correction ability: once an unsafe action is proposed, there is no mechanism to correct it. Thus, we propose CorrectionPlanne…

Reinforcement LearningAutonomous Driving

Off-Policy Correction For Multi-Agent Reinforcement Learning

2021-11-22 · Michał Zawalski, Błażej Osiński, Henryk Michalewski, Piotr Miłoś

Multi-agent reinforcement learning (MARL) provides a framework for problems involving multiple interacting agents. Despite apparent similarity to the single-agent case, multi-agent problems are often harder to train and …

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

FPGA-Based In-Vivo Calcium Image Decoding for Closed-Loop Feedback Applications

2022-12-09 · Zhe Chen, Garrett J. Blair, Chengdi Cao, Jim Zhou 외

Miniaturized calcium imaging is an emerging neural recording technique that has been widely used for monitoring neural activity on a large scale at a specific brain region of rats or mice. Most existing calcium-image ana…

TRACER: Training-Free Closed-Loop Structured Inference for Traffic Accident Reconstruction

2026-06-23 · Yanchen Guan, Chengyue Wang, Bin Rao, Haicheng Liao 외 arxiv

Traffic accident reconstruction is a forensic inverse problem that requires recovering physically consistent motion from sparse and heterogeneous evidence. Existing learning-based approaches predominantly optimize for se…

Closed-Loop Neural Operator-Based Observer of Traffic Density

2025-04-07 · Alice Harting, Karl Henrik Johansson, Matthieu Barreau

We consider the problem of traffic density estimation with sparse measurements from stationary roadside sensors. Our approach uses Fourier neural operators to learn macroscopic traffic flow dynamics from high-fidelity mi…

Density Estimation