paper-with-me

홈 › Papers

Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving

2025-06-13 · Boris Ivanovic, Cristiano Saltori, Yurong You, Yan Wang, Wenjie Luo, Marco Pavone

Autoregressive Transformers are increasingly being deployed as end-to-end robot and autonomous vehicle (AV) policy architectures, owing to their scalability and potential to leverage internet-scale pretraining for generalization. Accordingly, tokenizing sensor data efficiently is paramount to ensuring the real-time feasibility of such architectures on embedded hardware. To this end, we present an efficient triplane-based multi-camera tokenization strategy that leverages recent advances in 3D neural reconstruction and rendering to produce sensor tokens that are agnostic to the number of input cameras and their resolution, while explicitly accounting for their geometry around an AV. Experiments on a large-scale AV dataset and state-of-the-art neural simulator demonstrate that our approach yields significant savings over current image patch-based tokenization strategies, producing up to 72% fewer tokens, resulting in up to 50% faster policy inference while achieving the same open-loop motion planning accuracy and improved offroad rates in closed-loop driving simulations.

📄 PDF Abstract BibTeX arXiv:2506.12251

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Planning

Similar Papers 제목 키워드 기반

Freeplane: Unlocking Free Lunch in Triplane-Based Sparse-View Reconstruction Models

2024-06-02 · Wenqiang Sun, Zhengyi Wang, Shuo Chen, Yikai Wang 외

Creating 3D assets from single-view images is a complex task that demands a deep understanding of the world. Recently, feed-forward 3D generative models have made significant progress by training large reconstruction mod…

3D geometry

SYM3D: Learning Symmetric Triplanes for Better 3D-Awareness of GANs

2024-06-10 · Jing Yang, Kyle Fogarty, Fangcheng Zhong, Cengiz Oztireli

Despite the growing success of 3D-aware GANs, which can be trained on 2D images to generate high-quality 3D assets, they still rely on multi-view images with camera annotations to synthesize sufficient details from all v…

Text to 3D

DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding

2026-03-19 · Dong Zhuo, Wenzhao Zheng, Sicheng Zuo, Siming Yan 외 arxiv

With the growing adoption of vision-language-action models and world models in autonomous driving systems, scalable image tokenization becomes crucial as the interface for the visual modality. However, most existing toke…

Semantic SegmentationImage ReconstructionAutonomous Driving

Temporal Triplane Transformers as Occupancy World Models

2025-03-10 · Haoran Xu, Peixi Peng, Guang Tan, Yiqian Chang 외

World models aim to learn or construct representations of the environment that enable the prediction of future scenes, thereby supporting intelligent motion planning. However, existing models often struggle to produce fi…

Autonomous DrivingMotion Planning

RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models

2024-07-09 · BoWen Zhang, Yiji Cheng, Chunyu Wang, Ting Zhang 외

We present RodinHD, which can generate high-fidelity 3D avatars from a portrait image. Existing methods fail to capture intricate details such as hairstyles which we tackle in this paper. We first identify an overlooked …

DecoderScheduling