paper-with-me

홈 › Papers

GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation

2025-09-19 · Quanhao Qian, Guoyang Zhao, Gongjie Zhang, Jiuniu Wang, Ran Xu, Junlong Gao, Deli Zhao arxiv

Effective robotic manipulation relies on a precise understanding of 3D scene geometry, and one of the most straightforward ways to acquire such geometry is through multi-view observations. Motivated by this, we present GP3 -- a 3D geometry-aware robotic manipulation policy that leverages multi-view input. GP3 employs a spatial encoder to infer dense spatial features from RGB observations, which enable the estimation of depth and camera parameters, leading to a compact yet expressive 3D scene representation tailored for manipulation. This representation is fused with language instructions and translated into continuous actions via a lightweight policy head. Comprehensive experiments demonstrate that GP3 consistently outperforms state-of-the-art methods on simulated benchmarks. Furthermore, GP3 transfers effectively to real-world robots without depth sensors or pre-mapped environments, requiring only minimal fine-tuning. These results highlight GP3 as a practical, sensor-agnostic solution for geometry-aware robotic manipulation.

📄 PDF Abstract BibTeX arXiv:2509.15733

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

G-NeRF: Geometry-enhanced Novel View Synthesis from Single-View Images

2024-04-11 · CVPR 2024 1 · Zixiong Huang, Qi Chen, Libo Sun, Yifan Yang 외

Novel view synthesis aims to generate new view images of a given view image collection. Recent attempts address this problem relying on 3D geometry priors (e.g., shapes, sizes, and positions) learned from multi-view imag…

3D geometryNeRFNovel View Synthesis

Geometry-Aware Reference Synthesis for Multi-View Image Super-Resolution

2022-07-18 · Ri Cheng, Yuqi Sun, Bo Yan, Weimin Tan 외

Recent multi-view multimedia applications struggle between high-resolution (HR) visual experience and storage or bandwidth constraints. Therefore, this paper proposes a Multi-View Image Super-Resolution (MVISR) task. It …

Image Super-ResolutionSuper-ResolutionVideo Super-Resolution

Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis

2022-04-13 · CVPR 2022 1 · Xuanmeng Zhang, Zhedong Zheng, Daiheng Gao, Bang Zhang 외

3D-aware image synthesis aims to generate images of objects from multiple views by learning a 3D representation. However, one key challenge remains: existing approaches lack geometry constraints, hence usually fail to ge…

3D-Aware Image Synthesis3D geometryImage Generation

ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection

2023-08-17 · ICCV 2023 1 · Tao Tu, Shun-Po Chuang, Yu-Lun Liu, Cheng Sun 외

We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxe…

3D Object Detectionobject-detectionObject Detection

MoGaFace: Momentum-Guided and Texture-Aware Gaussian Avatars for Consistent Facial Geometry

2025-08-02 · Yujian Liu, Linlang Cao, Chuang Chen, Fanyu Geng 외 arxiv

Existing 3D head avatar reconstruction methods adopt a two-stage process, relying on tracked FLAME meshes derived from facial landmarks, followed by Gaussian-based rendering. However, misalignment between the estimated m…