paper-with-me

홈 › Papers

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

2026-05-11 · Hao Wang, Xiaobao Wei, Jingyang He, Chengyu Bai, Chun-Kai Fan, Jiajun Cao, Jintao Chen, Ying Li, Shanyu Rong, Ming Lu, Xiaozhu Ju, Jian Tang, Shanghang Zhang arxiv

Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image data without explicit 3D geometric supervision, resulting in representations that lack accurate spatial awareness. Existing implicit spatial grounding methods partially address this by aligning VLA features with those of 3D-aware foundation models, but they rely on empirical layer search and perform alignment on LLM-level visual tokens where spatial structure has already been entangled with linguistic semantics, limiting both generalizability and geometric interpretability. We propose VEGA (Visual Encoder Grounding Alignment), a simple yet effective framework that directly aligns the output of the VLA's visual encoder with spatially-aware features from DINOv2-FiT3D, a DINOv2 model fine-tuned with multi-view consistent 3D Gaussian Splatting supervision. By performing alignment at the visual encoder output level, VEGA grounds spatial awareness before any linguistic entanglement occurs, offering a more interpretable and principled alignment target. The alignment is implemented via a lightweight projector trained with a cosine similarity loss alongside the standard action prediction objective, and is discarded at inference time, introducing no additional computational overhead. Extensive experiments on simulation benchmark and real-world manipulation tasks demonstrate that VEGA consistently outperforms existing implicit spatial grounding baselines, establishing a new state-of-the-art among implicit spatial grounding methods for VLA models.

📄 PDF Abstract BibTeX arXiv:2605.10485

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC

2025-08-06 · Guanyu Hu, Dimitrios Kollias, Xinyu Yang arxiv

Multimodal Emotion Recognition in Conversations remains a challenging task due to the complex interplay of textual, acoustic and visual signals. While recent models have improved performance via advanced fusion strategie…

Multimodal Emotion Recognition

Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models

2026-08-02 · Myeongkyun Kang, Yanting Yang, Xiaoxiao Li arxiv

Fine-grained visual representations are essential for medical image analysis, particularly when diagnostically relevant evidence is subtle and spatially localized. Modern transformer-based medical vision encoders must th…

Visual Question AnsweringSelf-Supervised LearningRepresentation LearningPhrase Grounding

SCVIB: Editable State-Conditioned Visual Instance Binding forMulti-Turn Personalized Localization

2026-08-14 · Xiongtai Yang, Ziyan He, Tao Wang arxiv

We introduce editable state-conditioned visual instance binding, a multi-turn localization setting in which several support-defined instances are introduced across turns and protocol-defined state events determine the fi…

Visual Grounding

GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

2025-12-02 · Peirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin 외 arxiv

Recent advances in multimodal large language models(MLLMs) have led to remarkable progress in visual grounding, enabling fine-grained cross-modal alignment between textual queries and image regions. However, transferring…

Domain GeneralizationSpatial ReasoningVisual Grounding

VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering

2025-12-12 · Zihu Wang, Boxun Xu, Yuxuan Xia, Peng Li arxiv

Large vision-language models (LVLMs) exhibit impressive ability to jointly reason over visual and textual inputs. However, they often produce outputs that are linguistically fluent but factually inconsistent with the vis…