paper-with-me

Papers

Visual Implicit Geometry Transformer for Autonomous Driving

2026-02-05 · Arsenii Shirokov, Mikhail Kuznetsov, Danila Stepochkin, Egor Evdokimov, Daniil Glazkov, Nikolay Patakin, Anton Konushin, Dmitry Senushkin arxiv

We introduce the Visual Implicit Geometry Transformer (ViGT), an autonomous driving geometric model that estimates continuous 3D occupancy fields from surround-view camera rigs. ViGT represents a step towards foundational geometric models for autonomous driving, prioritizing scalability, architectural simplicity, and generalization across diverse sensor configurations. Our approach achieves this through a calibration-free architecture, enabling a single model to adapt to different sensor setups. Unlike general-purpose geometric foundational models that focus on pixel-aligned predictions, ViGT estimates a continuous 3D occupancy field in a birds-eye-view (BEV) addressing domain-specific requirements. ViGT naturally infers geometry from multiple camera views into a single metric coordinate frame, providing a common representation for multiple geometric tasks. Unlike most existing occupancy models, we adopt a self-supervised training procedure that leverages synchronized image-LiDAR pairs, eliminating the need for costly manual annotations. We validate the scalability and generalizability of our approach by training our model on a mixture of five large-scale autonomous driving datasets (NuScenes, Waymo, NuPlan, ONCE, and Argoverse) and achieving state-of-the-art performance on the pointmap estimation task, with the best average rank across all evaluated baselines. We further evaluate ViGT on the Occ3D-nuScenes benchmark, where ViGT achieves comparable performance with supervised methods. The source code is publicly available at \href{https://github.com/whesense/ViGT}{https://github.com/whesense/ViGT}.

📄 PDF Abstract BibTeX arXiv:2602.05573

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

DVGT: Driving Visual Geometry Transformer

2025-12-18 · Sicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu 외 arxiv

Perceiving and reconstructing 3D scene geometry from visual inputs is crucial for autonomous driving. However, there still lacks a driving-targeted dense geometry perception model that can adapt to different scenarios an…

Autonomous Driving

GeoWorldAD: Geometry World Action Model for Autonomous Driving

2026-07-20 · Songyan Zhang, Jinyuan Tian, Hanbing Li, Daqi Liu 외 arxiv

Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances …

Collision AvoidanceTrajectory PlanningAutonomous Driving

DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale

2026-04-01 · Sicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu 외 arxiv

End-to-end autonomous driving has evolved from the conventional paradigm based on sparse perception into vision-language-action (VLA) models, which focus on learning language descriptions as an auxiliary task to facilita…

Trajectory PlanningAutonomous Driving

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

2026-08-24 · Yiren Lu, Xin Ye, Jiaming Liu, Jin Yao 외 arxiv

World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by co…

Trajectory PredictionAutonomous DrivingPoint Clouds

LiAuto-GeoX: Efficient Grounded Driving Transformer

2026-06-04 · Jiawei Lian, Haoyi Sun, Yang Wu, Lifu Mu 외 arxiv

Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an open challenge. Existing large-scale visual…

Trajectory PredictionScene UnderstandingAutonomous Driving3D Reconstruction