paper-with-me

홈 › Papers

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning

2026-06-01 · Huayi Zhou, Wei Gao, Dekun Lu, Ruiji Liu, Zhanqi Zhang, Ziyang Zhang, Jian Chen, Wenlve Zhou, Sheng Xu, Shumin Li, Kangyi Guo, Shichen Xu, Zixin Huang, Yongyi Su, Kui Jia arxiv

End-to-end manipulation policies, combined with web-scale pretrained Vision-Language Models (VLMs), show the promise for generalizable and dexterous robotic manipulation. However, they inherit two key limitations from 2D foundation models: 1) the reliance on 2D RGB inputs that ignores the intrinsically 3D nature of manipulation; and 2) the lack of spatial 3D alignment between input-output spaces as well as across diverse robot embodiments, camera setups, and trajectory datasets. In this paper, we present a series of contributions to address these issues. First, we introduce aligned vertex map and vertex spectrum -- a pixel-wise 3D representation that elevates 2D visual inputs to 3D, using camera calibration and optional depth. This novel input representation marries 3D awareness with the generalization of 2D large VLMs. Then, we propose to align the inputs and outputs of manipulation policies by expressing per-pixel 3D information of each camera view and robot actions to a shared coordinate. Based on this, we designate a canonical Bird's-Eye-View (BEV) alignment frame and innovatively propose to construct BEV images, producing a view-invariant representation robust to camera pose variations. To enable training and evaluation at scale, we develop a comprehensive data processing pipeline to perform such alignments; we also introduce a novel temporal alignment scheme for trajectories across diverse robots, human operators, and datasets. These contributions collectively mitigate input and output spatial-temporal misalignments, improving the consistency and generalization for real-world manipulation. Pretrained checkpoint, source code and data processing pipeline are available in https://hnuzhy.github.io/projects/Dex-BEV.

📄 PDF Abstract BibTeX arXiv:2606.02274

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

World Models for Learning Dexterous Hand-Object Interactions from Human Videos

2025-12-15 · Raktim Gautam Goswami, Amir Bar, David Fan, Tsung-Yen Yang 외 arxiv

Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While recent world models address interaction mode…

Generalizable Domain Adaptation for Sim-and-Real Policy Co-Training

2025-09-23 · Shuo Cheng, Liqian Ma, Zhenyang Chen, Ajay Mandlekar 외 arxiv

Behavior cloning has shown promise for robot manipulation, but real-world demonstrations are costly to acquire at scale. While simulated data offers a scalable alternative, particularly with advances in automated demonst…

Robot ManipulationDomain Adaptation

Enhancing Tactile-based Reinforcement Learning for Robotic Control

2025-10-24 · Elle Miller, Trevor McInroe, David Abel, Oisin Mac Aodha 외 arxiv

Achieving safe, reliable real-world robotic manipulation requires agents to evolve beyond vision and incorporate tactile sensing to overcome sensory deficits and reliance on idealised state information. Despite its poten…

Self-Supervised LearningReinforcement Learning

TypeTele: Releasing Dexterity in Teleoperation by Dexterous Manipulation Types

2025-07-02 · Yuhao Lin, Yi-Lin Wei, Haoran Liao, Mu Lin 외 arxiv

Dexterous teleoperation plays a crucial role in robotic manipulation for real-world data collection and remote robot control. Previous dexterous teleoperation mostly relies on hand retargeting to closely mimic human hand…

Sumo: Dynamic and Generalizable Whole-Body Loco-Manipulation

2026-04-09 · John Z. Zhang, Maks Sorokin, Jan Brüdigam, Brandon Hung 외 arxiv

This paper presents a sim-to-real approach that enables legged robots to dynamically manipulate large and heavy objects with whole-body dexterity. Our key insight is that by performing test-time steering of a pre-trained…