paper-with-me

홈 › Papers

Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning

2026-06-01 · Ziyang Yao, Zeyu Zhu, YunCheng Jiang, Zibin Guo, Huijing Zhao arxiv

Discrete visual tokens should provide a compact representation for both token-based world modeling and planning in autonomous driving. However, most tokenizers are inherited from image generation and are optimized mainly for pixel reconstruction, which may leave a gap between what is easy to generate and what is useful to decode for driving decisions. We present a representation-guided and geometry-enhanced tokenizer that learns discrete tokens under joint supervision. The tokenizer aligns its discrete bottleneck with a frozen DINO feature space through feature decoding, while preserving appearance via RGB reconstruction with perceptual and adversarial losses. To inject geometric state-related cues, we add adjacent-frame depth and relative-pose supervision during training and stabilize joint objectives with multi-codebook quantization. We evaluate the same learned tokens with a lightweight planning readout and a GPT-style next-token world model. Experiments on NAVSIM show improved reconstruction fidelity and representation consistency, competitive planning performance under a fixed decoder, and better generative quality under matched settings.

📄 PDF Abstract BibTeX arXiv:2606.01935

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingImage Generation

Similar Papers 제목 키워드 기반

USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

2026-08-19 · Li-Heng Chen, Haokai Pang, Chengye Su, Jiarun Liu 외 arxiv

Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians)…

Representation LearningDynamic ReconstructionScene UnderstandingAutonomous Driving

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

2026-08-02 · Ruiteng Zhao, Zhengshen Zhang, Yue Su, Wenshuo Wang 외 hf

World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficie…

Geometry-Grounded Unified 3D Perception for Autonomous Driving

2026-08-13 · Longfei Xu, Xiaohui Wang, Zehao Huang, Han Li 외 arxiv

Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pr…

3D Object DetectionAutonomous DrivingDepth Estimation

UniDWM: Towards a Unified Driving World Model via Multifaceted Representation Learning

2026-02-02 · Shuai Liu, Siheng Ren, Xiaoyao Zhu, Quanmin Liang 외 arxiv

Achieving reliable and efficient planning in complex driving environments requires a model that can reason over the scene's geometry, appearance, and dynamics. We present UniDWM, a unified driving world model that advanc…

Representation LearningTrajectory PlanningAutonomous Driving

Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

2026-02-25 · Matthew Strong, Wei-Jer Chang, Quentin Herau, Jiezhi Yang 외 arxiv

Ego-centric driving videos available online provide an abundant source of visual data for autonomous driving, yet their lack of annotations makes it difficult to learn representations that capture both semantic structure…

Semantic SegmentationAutonomous Driving