paper-with-me

홈 › Papers

Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning

2025-10-04 · Sunghwan Kim, Woojeh Chung, Zhirui Dai, Dwait Bhatt, Arth Shukla, Hao Su, Yulun Tian, Nikolay Atanasov arxiv

In this paper, we demonstrate that mobile manipulation policies utilizing a 3D latent map achieve stronger spatial and temporal reasoning than policies relying solely on images. We introduce Seeing the Bigger Picture (SBP), an end-to-end policy learning approach that operates directly on a 3D map of latent features. In SBP, the map extends perception beyond the robot's current field of view and aggregates observations over long horizons. Our mapping approach incrementally fuses multiview observations into a grid of scene-specific latent features. A pre-trained, scene-agnostic decoder reconstructs target embeddings from these features and enables online optimization of the map features during task execution. A policy, trainable with behavior cloning or reinforcement learning, treats the latent map as a state variable and uses global context from the map obtained via a 3D feature aggregator. We evaluate SBP on scene-level mobile manipulation and sequential tabletop manipulation tasks. Our experiments demonstrate that SBP (i) reasons globally over the scene, (ii) leverages the map as long-horizon memory, and (iii) outperforms image-based policies in both in-distribution and novel scenes, e.g., improving the success rate by 15% for the sequential manipulation task.

📄 PDF Abstract BibTeX arXiv:2510.03885

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

More Words and Bigger Pictures

2013-06-01 · SEMEVAL 2013 6 · David Forsyth
Object Recognition

50 Ways to Bake a Cookie: Mapping the Landscape of Procedural Texts

2022-10-31 · Moran Mizrahi, Dafna Shahaf

The web is full of guidance on a wide variety of tasks, from changing the oil in your car to baking an apple pie. However, as content is created independently, a single task could have thousands of corresponding procedur…

CLIPTER: Looking at the Bigger Picture in Scene Text Recognition

2023-01-18 · ICCV 2023 1 · Aviad Aberdam, David Bensaïd, Alona Golts, Roy Ganz 외

Reading text in real-world scenarios often requires understanding the context surrounding it, especially when dealing with poor-quality text. However, current scene text recognizers are unaware of the bigger picture as t…

Language ModelingLanguage ModellingScene Text Recognition

Reduction of Class Activation Uncertainty with Background Information

2023-05-05 · H M Dipu Kabir

Multitask learning is a popular approach to training high-performing neural networks with improved generalization. In this paper, we propose a background class to achieve improved generalization at a lower computation co…

ClassificationFine-Grained Image ClassificationImage ClassificationSatellite Image Classification

Image Semantic Transformation: Faster, Lighter and Stronger

2018-03-27 · Dasong Li, Jianbo Wang

We propose Image-Semantic-Transformation-Reconstruction-Circle(ISTRC) model, a novel and powerful method using facenet's Euclidean latent space to understand the images. As the name suggests, ISTRC construct the circle, …