paper-with-me

Papers

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation

2025-05-15 · Zibin Dong, Fei Ni, Yifu Yuan, Yinchuan Li, Jianye Hao

We present EmbodiedMAE, a unified 3D multi-modal representation for robot manipulation. Current approaches suffer from significant domain gaps between training datasets and robot manipulation tasks, while also lacking model architectures that can effectively incorporate 3D information. To overcome these limitations, we enhance the DROID dataset with high-quality depth maps and point clouds, constructing DROID-3D as a valuable supplement for 3D embodied vision research. Then we develop EmbodiedMAE, a multi-modal masked autoencoder that simultaneously learns representations across RGB, depth, and point cloud modalities through stochastic masking and cross-modal fusion. Trained on DROID-3D, EmbodiedMAE consistently outperforms state-of-the-art vision foundation models (VFMs) in both training efficiency and final performance across 70 simulation tasks and 20 real-world robot manipulation tasks on two robot platforms. The model exhibits strong scaling behavior with size and promotes effective policy learning from 3D inputs. Experimental results establish EmbodiedMAE as a reliable unified 3D multi-modal VFM for embodied AI systems, particularly in precise tabletop manipulation settings where spatial perception is critical.

📄 PDF Abstract BibTeX arXiv:2505.10105

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

An Attention-based Model for Robust Forecasting with Missing Modality

2026-06-11 · Zhitian Zhang, Wenjie Zi, Yunduz Rakhmangulova, Saghar Irandoust 외 arxiv

Learning with missing modalities is a fundamental challenge in multimodal robot learning, as real-world robotic systems often operate in environments with incomplete sensor data. Attention-based models are appealing for …

Trajectory PredictionRobot Manipulation

Autonomous Laparoscope Control through Unified Mechanics-Based Representation of Multimodal Intraoperative Information

2026-05-06 · Xiaojian Li, Jin Fang, Yudong Shi, Xilin Xiao 외 arxiv

Laparoscope-holding robots can provide surgeons with a stable laparoscopic field of view (FOV) and reduce the burden on human assistants. To maintain an ideal intraoperative FOV, the robot must continuously adjust the la…

Visual Tracking

Audio Visual Language Maps for Robot Navigation

2023-03-13 · Chenguang Huang, Oier Mees, Andy Zeng, Wolfram Burgard

While interacting in the world is a multi-sensory experience, many robots continue to predominantly rely on visual perception to map and navigate in their environments. In this work, we propose Audio-Visual-Language Maps…

NavigateRobot Navigation

AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

2025-02-15 · Ruoxuan Feng, Jiangyu Hu, Wenke Xia, Tianci Gao 외

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into rob…

Representation LearningTransfer Learning

MEM: Multi-Modal Elevation Mapping for Robotics and Learning

2023-09-28 · Gian Erni, Jonas Frey, Takahiro Miki, Matias Mattamala 외

Elevation maps are commonly used to represent the environment of mobile robots and are instrumental for locomotion and navigation tasks. However, pure geometric information is insufficient for many field applications tha…

ColorizationGPUHuman DetectionLine Detection