paper-with-me

홈 › Papers

VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

2026-06-10 · Jin Yao, Dhruva Dixith Kurra, Tom Lampo, Zezhou Cheng, Danhua Guo, Burhan Yaman arxiv

Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal. We introduce VLGA, the first vision-language-action model supervised to reconstruct the dense 3D world it drives through. VLGA introduces geometry as a fourth modality alongside vision, language, and action through a dedicated expert supervised by a per-pixel pointmap regression loss against LiDAR. Extensive experiments conducted on challenging nuScenes and Bench2Drive datasets for open-loop and closed-loop evaluations, respectively, show the superiority of VLGA over counterpart VLA methods. In particular, on open-loop nuScenes, VLGA sets a new state of the art among VLA methods without ego status, with the lowest L2 (0.50\,m average) and 3-second collision rate (0.18\%). On closed-loop Bench2Drive, VLGA attains the state-of-the-art driving score of 79.08, +0.71 over the strongest prior VLA, at comparable efficiency and comfort.

📄 PDF Abstract BibTeX arXiv:2606.12396

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale

2026-04-01 · Sicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu 외 arxiv

End-to-end autonomous driving has evolved from the conventional paradigm based on sparse perception into vision-language-action (VLA) models, which focus on learning language descriptions as an auxiliary task to facilita…

Trajectory PlanningAutonomous Driving

GeoWorldAD: Geometry World Action Model for Autonomous Driving

2026-07-20 · Songyan Zhang, Jinyuan Tian, Hanbing Li, Daqi Liu 외 arxiv

Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances …

Collision AvoidanceTrajectory PlanningAutonomous Driving

Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency Relationships

2022-03-27 · CVPR 2022 1 · Chao Lou, Wenjuan Han, Yuhuan Lin, Zilong Zheng

Understanding realistic visual scene images together with language descriptions is a fundamental task towards generic visual understanding. Previous works have shown compelling comprehensive results by building hierarchi…

Contrastive LearningPhrase Grounding

Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion

2026-02-24 · Jiaru Zhang, Manav Gagvani, Can Cui, Juntong Peng 외 arxiv

Large Language Models (LLMs) and Vision-Language Models (VLMs) have emerged as promising candidates for end-to-end autonomous driving. However, these models typically face challenges in inference latency, action precisio…

Autonomous Driving

sensVLA: Spatially-Grounded Vision-Language-Action Model for Autonomous Wheel Loader

2026-09-15 · Gopi Krishna Erabati, Bjarne Johannsen, Angus Stewart, Vardeep Singh Sandhu arxiv

Autonomous wheel-loader control requires joint reasoning over task semantics, egocentric vision, proprioception, and 3D scene geometry. We present sensVLA, a Vision-Language-Action (VLA) architecture that combines a Qwen…