paper-with-me

홈 › Papers

G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models

2026-06-23 · Yue Peng, Yongzhe Zhao, Artur Habuda, Khuyen Pham, Yanheng Zhu, Tran Nguyen Le, Fares Abu-Dakka, Li Guo arxiv

Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their visual tokens remain grounded in 2D image coordinates rather than the calibrated geometry of the robot's cameras -- a mismatch especially pronounced in multi-camera setups, where views are coupled by known intrinsics and extrinsics yet processed as independent images. We propose G$^3$VLA, a camera-aware geometric module that injects calibrated structure into the visual-token stream of a pretrained VLA without altering its action space or imitation objective, combining intrinsic-conditioned ray embeddings, projective positional encoding (PRoPE), and bidirectional cross-view fusion. Geometric supervision is provided either from ground-truth point maps when available, or from confidence-gated $π^3$X teacher predictions, requiring no depth sensors or manual annotations. Instantiated on $π_0$, G$^3$VLA yields consistent gains across the LIBERO suites, RoboCasa24, RoboTwin2.0, and real-robot settings, with the largest improvements on spatially and object-sensitive tasks. We further validate on $π_{0.5}$ and GR00T 1.5, with results suggesting that geometric transfer is most effective when geometry-aware tokens have direct access to the action generation pathway. Our project page is at https://sites.google.com/view/g3vla

📄 PDF Abstract BibTeX arXiv:2606.24472

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models

2025-09-18 · Haobo Yang, Minghao Guo, Dequan Yang, Wenyu Wang arxiv

Contemporary deep learning models have achieved impressive performance in image classification by primarily leveraging statistical regularities within large datasets, but they rarely incorporate structured insights drawn…

Image Classification

Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings

2026-02-05 · Grégoire Dhimoïla, Thomas Fel, Victor Boutin, Agustin Picard arxiv

Vision-language models (VLMs) align images and text with remarkable success, yet the geometry of their shared embedding space remains poorly understood. To probe this geometry, we begin from the Iso-Energy Assumption, wh…

Input-level Inductive Biases for 3D Reconstruction

2021-12-06 · CVPR 2022 1 · Wang Yifan, Carl Doersch, Relja Arandjelović, João Carreira 외

Much of the recent progress in 3D vision has been driven by the development of specialized architectures that incorporate geometrical inductive biases. In this paper we tackle 3D reconstruction using a domain agnostic ar…

3D ReconstructionDepth Estimation

SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning

2026-04-23 · Chan Yeong Hwang, Miso Choi, Sunghyun On, Jinkyu Kim 외 arxiv

Understanding visual scenes requires not only recognizing objects but also reasoning about their spatial relationships. Unlike general vision-language tasks, spatial reasoning requires integrating multiple inductive bias…

Spatial Reasoning

Injecting structural hints: Using language models to study inductive biases in language learning

2023-04-25 · Isabel Papadimitriou, Dan Jurafsky

Both humans and large language models are able to learn language without explicit structural supervision. What inductive biases make this learning possible? We address this fundamental cognitive question by leveraging tr…

Inductive BiasTransfer Learning