paper-with-me

Papers

SEM: Enhancing Spatial Understanding for Robust Robot Manipulation

2025-05-22 · Xuewu Lin, Tianwei Lin, Lichao Huang, Hongyu Xie, Yiwei Jin, Keyu Li, Zhizhong Su

A key challenge in robot manipulation lies in developing policy models with strong spatial understanding, the ability to reason about 3D geometry, object relations, and robot embodiment. Existing methods often fall short: 3D point cloud models lack semantic abstraction, while 2D image encoders struggle with spatial reasoning. To address this, we propose SEM (Spatial Enhanced Manipulation model), a novel diffusion-based policy framework that explicitly enhances spatial understanding from two complementary perspectives. A spatial enhancer augments visual representations with 3D geometric context, while a robot state encoder captures embodiment-aware structure through graphbased modeling of joint dependencies. By integrating these modules, SEM significantly improves spatial understanding, leading to robust and generalizable manipulation across diverse tasks that outperform existing baselines.

📄 PDF Abstract BibTeX arXiv:2505.16196

Code (0)

등록된 구현이 없습니다.

Tasks

3D geometryRobot ManipulationSpatial Reasoning

Similar Papers 제목 키워드 기반

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum

2026-05-20 · Zhizhao Liang, Yi-Lin Wei, Xuhang Chen, Mu Lin 외 arxiv

In this paper, we explore spatial-aware humanoid whole-body manipulation task. Compared with tabletop settings, this task poses two key challenges: 1) Spatial understanding is challenging in complex 3D environments with …

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding

2026-05-28 · Zhongyu Xia, Yousen Tang, Bingqing Wei, Yongtao Wang arxiv

Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficiency manifests as three intertwined challe…

Scene Understanding

Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy

2025-11-14 · Vinit Mehta, Charu Sharma, Karthick Thiyagarajan arxiv

With the rapid advancement of artificial intelligence and robotics, the integration of Large Language Models (LLMs) with 3D vision is emerging as a transformative approach to enhancing robotic sensing technologies. This …

Scene Understanding3D Generation

ManiBox: Enhancing Spatial Grasping Generalization via Scalable Simulation Data Generation

2024-11-04 · Hengkai Tan, Xuezhou Xu, Chengyang Ying, Xinyi Mao 외

Learning a precise robotic grasping policy is crucial for embodied agents operating in complex real-world manipulation tasks. Despite significant advancements, most models still struggle with accurate spatial positioning…

Robotic Grasping

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning

2025-10-16 · Jinrui Liu, Bingyan Nie, Boyu Li, Yaran Chen 외 arxiv

Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully. Despite the success of large language models and vision l…

Reinforcement LearningRobot Task Planning