paper-with-me

홈 › Papers

When Learning Is Out of Reach, Reset: Generalization in Autonomous Visuomotor Reinforcement Learning

2023-03-30 · Zichen Zhang, Luca Weihs

Episodic training, where an agent's environment is reset after every success or failure, is the de facto standard when training embodied reinforcement learning (RL) agents. The underlying assumption that the environment can be easily reset is limiting both practically, as resets generally require human effort in the real world and can be computationally expensive in simulation, and philosophically, as we'd expect intelligent agents to be able to continuously learn without intervention. Work in learning without any resets, i.e{.} Reset-Free RL (RF-RL), is promising but is plagued by the problem of irreversible transitions (e.g{.} an object breaking) which halt learning. Moreover, the limited state diversity and instrument setup encountered during RF-RL means that works studying RF-RL largely do not require their models to generalize to new environments. In this work, we instead look to minimize, rather than completely eliminate, resets while building visual agents that can meaningfully generalize. As studying generalization has previously not been a focus of benchmarks designed for RF-RL, we propose a new Stretch Pick-and-Place benchmark designed for evaluating generalizations across goals, cosmetic variations, and structural changes. Moreover, towards building performant reset-minimizing RL agents, we propose unsupervised metrics to detect irreversible transitions and a single-policy training mechanism to enable generalization. Our proposed approach significantly outperforms prior episodic, reset-free, and reset-minimizing approaches achieving higher success rates with fewer resets in Stretch-P\&P and another popular RF-RL benchmark. Finally, we find that our proposed approach can dramatically reduce the number of resets required for training other embodied tasks, in particular for RoboTHOR ObjectNav we obtain higher success rates than episodic approaches using 99.97\% fewer resets.

📄 PDF Abstract BibTeX arXiv:2303.17600

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Learning to Autonomously Reach Objects with NICO and Grow-When-Required Networks

2022-10-14 · Nima Rahrakhshan, Matthias Kerzel, Philipp Allgeuer, Nicolas Duczek 외

The act of reaching for an object is a fundamental yet complex skill for a robotic agent, requiring a high degree of visuomotor control and coordination. In consideration of dynamic environments, a robot capable of auton…

Object

Prepare Before You Act: Learning From Humans to Rearrange Initial States

2025-09-22 · Yinlong Dai, Andre Keyser, Dylan P. Losey arxiv

Imitation learning (IL) has proven effective across a wide range of manipulation tasks. However, IL policies often struggle when faced with out-of-distribution observations; for instance, when the target object is in a p…

Unsupervised Visuomotor Control through Distributional Planning Networks

2019-02-14 · Tianhe Yu, Gleb Shevchuk, Dorsa Sadigh, Chelsea Finn

While reinforcement learning (RL) has the potential to enable robots to autonomously acquire a wide range of skills, in practice, RL usually requires manual, per-task engineering of reward functions, especially in real w…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning

2026-03-16 · Patrick Yin, Tyler Westenbroek, Zhengyu Zhang, Joshua Tran 외 arxiv

Reinforcement learning in massively parallel physics simulations has driven major progress in sim-to-real robot learning. However, current approaches remain brittle and task-specific, relying on extensive per-task engine…

Reinforcement Learning

ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning

2025-12-18 · Zihan Zhou, Animesh Garg, Ajay Mandlekar, Caelan Garrett arxiv

Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines task decomposition, data generation, imitation learning, and motion planning to form…

Reinforcement LearningMotion Planning