paper-with-me

Papers

Rethinking Implicit Spatial Representation in Visuomotor Policy Learning

2026-06-13 · Xiangyu Chen, Yuxuan Hu, Chuhao Zhou, Jianfei Yang arxiv

Generative model-based imitation learning has become a widely adopted paradigm for robotic manipulation, where policy performance depends critically on the conditioned visual representations. Although spatial softmax-based representations have been adopted in prior visuomotor policies, their effectiveness and underlying mechanisms remain insufficiently understood. This work rethinks the use of spatial softmax pooling: do such implicit spatial representations provide effective and stable visual features for robotic manipulation? Through systematic studies of different pooling methods in visual encoders, we find that this pooling operation produces compact and stable spatial representations, which outperform feature-value representations, despite using substantially fewer dimensions. Complementary saliency analysis further suggests that these spatial representations guide the encoder to focus more consistently on task-relevant regions. However, this advantage is limited by a representation bottleneck in current visual encoders: repeated downsampling operations weaken fine-grained spatial information before the action-generation module can use it, especially under low-resolution observations. Motivated by these findings, we propose PRISM, a visual encoder that preserves multiscale implicit spatial information through top-down cross-attention fusion. Experiments across multiple tasks and policy backbones show consistent improvements. In particular, on the low-resolution, high-precision ToolHang task, PRISM shows clear gains, improving the average success rate from 5.0% to 13.4% while increasing parameters by only 15.4%. These results support the use of multiscale implicit spatial representations as an effective and efficient design principle for robotic manipulation.

📄 PDF Abstract BibTeX arXiv:2606.15232

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies

2025-08-11 · Yinpei Dai, Jayjun Lee, Yichi Zhang, Ziqiao Ma 외 arxiv

In this paper, we propose AimBot, a lightweight visual augmentation technique that provides explicit spatial cues to improve visuomotor policy learning in robotic manipulation. AimBot overlays shooting lines and scope re…

StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception

2026-05-11 · Evans Han, Yunfan Jiang, Yingke Wang, Haoyue Xiao 외 arxiv

Recent advances in robot imitation learning have produced powerful visuomotor policies that manipulate diverse objects from visual inputs. However, monocular observations lack depth information, which is critical for pre…

Point Clouds

Out-of-Distribution Recovery with Object-Centric Keypoint Inverse Policy for Visuomotor Imitation Learning

2024-11-05 · George Jiayuan Gao, Tianyu Li, Nadia Figueroa

We propose an object-centric recovery (OCR) framework to address the challenges of out-of-distribution (OOD) scenarios in visuomotor policy learning. Previous behavior cloning (BC) methods rely heavily on a large amount …

Continual LearningImitation LearningObjectOptical Character Recognition (OCR)

Introspective Visuomotor Control: Exploiting Uncertainty in Deep Visuomotor Control for Failure Recovery

2021-03-22 · Chia-Man Hung, Li Sun, Yizhe Wu, Ioannis Havoutis 외

End-to-end visuomotor control is emerging as a compelling solution for robot manipulation tasks. However, imitation learning-based visuomotor control approaches tend to suffer from a common limitation, lacking the abilit…

Imitation LearningRobot Manipulation

Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning

2025-08-21 · Yijun Liu, Yuwei Liu, Yuan Meng, Jieheng Zhang 외 arxiv

Vision-centric hierarchical embodied models have demonstrated strong potential. However, existing methods lack spatial awareness capabilities, limiting their effectiveness in bridging visual plans to actionable control i…

Spatial ReasoningVideo Generation