paper-with-me

Papers

Efficient Training of Generalizable Visuomotor Policies via Control-Aware Augmentation

2024-01-17 · Yinuo Zhao, Kun Wu, Tianjiao Yi, Zhiyuan Xu, Xiaozhu Ju, Zhengping Che, Chi Harold Liu, Jian Tang

Improving generalization is one key challenge in embodied AI, where obtaining large-scale datasets across diverse scenarios is costly. Traditional weak augmentations, such as cropping and flipping, are insufficient for improving a model's performance in new environments. Existing data augmentation methods often disrupt task-relevant information in images, potentially degrading performance. To overcome these challenges, we introduce EAGLE, an efficient training framework for generalizable visuomotor policies that improves upon existing methods by (1) enhancing generalization by applying augmentation only to control-related regions identified through a self-supervised control-aware mask and (2) improving training stability and efficiency by distilling knowledge from an expert to a visuomotor student policy, which is then deployed to unseen environments without further fine-tuning. Comprehensive experiments on three domains, including the DMControl Generalization Benchmark, the enhanced Robot Manipulation Distraction Benchmark, and a long-sequential drawer-opening task, validate the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2401.09258

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationReinforcement Learning (RL)Robot Manipulation

Similar Papers 제목 키워드 기반

UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies

2025-10-02 · Harsh Gupta, Xiaofeng Guo, Huy Ha, Chuer Pan 외 arxiv

We introduce UMI-on-Air, a framework for embodiment-aware deployment of embodiment-agnostic manipulation policies. Our approach leverages diverse, unconstrained human demonstrations collected with a handheld gripper (UMI…

Adversarial Feature Training for Generalizable Robotic Visuomotor Control

2019-09-17 · Xi Chen, Ali Ghadirzadeh, Mårten Björkman, Patric Jensfelt

Deep reinforcement learning (RL) has enabled training action-selection policies, end-to-end, by learning a function which maps image pixels to action outputs. However, it's application to visuomotor robotic policy traini…

Deep Reinforcement LearningReinforcement LearningReinforcement Learning (RL)Transfer Learning

PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model

2025-11-03 · Wenqi Liang, Gan Sun, Yao He, Jiahua Dong 외 arxiv

Vision-Language-Action models (VLAs) are emerging as powerful tools for learning generalizable visuomotor control policies. However, current VLAs are mostly trained on large-scale image-text-action data and remain limite…

Scene Understanding

A Latency-Aware Framework for Visuomotor Policy Learning on Industrial Robots

2026-02-15 · Daniel Ruan, Salma Mozaffari, Sigrid Adriaenssens, Arash Adel arxiv

Industrial robots are increasingly deployed in contact-rich construction and manufacturing tasks that involve uncertainty and long-horizon execution. While learning-based visuomotor policies offer a promising alternative…

Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents

2025-07-31 · Shaofei Cai, Zhancun Mu, Haiwen Xia, Bowei Zhang 외 arxiv

While Reinforcement Learning (RL) has achieved remarkable success in language modeling, its triumph hasn't yet fully translated to visuomotor agents. A primary challenge in RL models is their tendency to overfit specific…

Zero-shot GeneralizationReinforcement LearningSpatial Reasoning