paper-with-me

홈 › Papers

Real-world Reinforcement Learning from Suboptimal Interventions

2025-12-30 · Yinuo Zhao, Huiqian Jin, Lechun Jiang, Xinyi Zhang, Kun Wu, Pei Ren, Zhiyuan Xu, Zhengping Che, Lei Sun, Dapeng Wu, Chi Harold Liu, Jian Tang arxiv

Real-world reinforcement learning (RL) offers a promising approach to training precise and dexterous robotic manipulation policies in an online manner, enabling robots to learn from their own experience while gradually reducing human labor. However, prior real-world RL methods often assume that human interventions are optimal across the entire state space, overlooking the fact that even expert operators cannot consistently provide optimal actions in all states or completely avoid mistakes. Indiscriminately mixing intervention data with robot-collected data inherits the sample inefficiency of RL, while purely imitating intervention data can ultimately degrade the final performance achievable by RL. The question of how to leverage potentially suboptimal and noisy human interventions to accelerate learning without being constrained by them thus remains open. To address this challenge, we propose SiLRI, a state-wise Lagrangian reinforcement learning algorithm for real-world robot manipulation tasks. Specifically, we formulate the online manipulation problem as a constrained RL optimization, where the constraint bound at each state is determined by the uncertainty of human interventions. We then introduce a state-wise Lagrange multiplier and solve the problem via a min-max optimization, jointly optimizing the policy and the Lagrange multiplier to reach a saddle point. Built upon a human-as-copilot teleoperation system, our algorithm is evaluated through real-world experiments on diverse manipulation tasks. Experimental results show that SiLRI effectively exploits human suboptimal interventions, reducing the time required to reach a 90% success rate by at least 50% compared with the state-of-the-art RL method HIL-SERL, and achieving a 100% success rate on long-horizon manipulation tasks where other RL methods struggle to succeed. Project website: https://silri-rl.github.io/.

📄 PDF Abstract BibTeX arXiv:2512.24288

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningRobot Manipulation

Similar Papers 제목 키워드 기반

ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning

2026-06-15 · Wei Xiao, Weiliang Tang, Yuying Ge, Hui Zhou 외 arxiv

Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models. However, enabling seamless humanoid interventions is a formidable systems challenge due to complex whole-body …

Reinforcement Learning

Prescriptive Process Monitoring Under Resource Constraints: A Reinforcement Learning Approach

2023-07-13 · Mahmoud Shoush, Marlon Dumas

Prescriptive process monitoring methods seek to optimize the performance of business processes by triggering interventions at runtime, thereby increasing the probability of positive case outcomes. These interventions are…

Conformal Predictionreinforcement-learningReinforcement Learning

OHP-RL: Online Human Preference as Guidance in Reinforcement Learning for Robot Manipulation

2026-05-15 · Yunyang Mo, Jian Li, Qiwei Wu, Yihang Kang 외 arxiv

While reinforcement learning (RL) enables robots to acquire skills autonomously, its real-world deployment is severely limited by inefficient and unsafe exploration. Human-in-the-loop interventions offer a practical solu…

Reinforcement LearningRobot Manipulation

SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations

2025-07-11 · Peter Crowley, Zachary Serlin, Tyler Paine, Makai Mann 외 arxiv

Inverse Reinforcement Learning (IRL) presents a powerful paradigm for learning complex robotic tasks from human demonstrations. However, most approaches make the assumption that expert demonstrations are available, which…

Reinforcement Learning

Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems

2020-08-30 · Vinicius G. Goecks

Recent successes combine reinforcement learning algorithms and deep neural networks, despite reinforcement learning not being widely applied to robotics and real world scenarios. This can be attributed to the fact that c…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)