paper-with-me

홈 › Papers

EMMA: End-to-End Multimodal Model for Autonomous Driving

2024-10-30 · Jyh-Jing Hwang, Runsheng Xu, Hubert Lin, Wei-Chih Hung, Jingwei Ji, Kristy Choi, Di Huang, Tong He, Paul Covington, Benjamin Sapp, Yin Zhou, James Guo, Dragomir Anguelov, Mingxing Tan

We introduce EMMA, an End-to-end Multimodal Model for Autonomous driving. Built on a multi-modal large language model foundation, EMMA directly maps raw camera sensor data into various driving-specific outputs, including planner trajectories, perception objects, and road graph elements. EMMA maximizes the utility of world knowledge from the pre-trained large language models, by representing all non-sensor inputs (e.g. navigation instructions and ego vehicle status) and outputs (e.g. trajectories and 3D locations) as natural language text. This approach allows EMMA to jointly process various driving tasks in a unified language space, and generate the outputs for each task using task-specific prompts. Empirically, we demonstrate EMMA's effectiveness by achieving state-of-the-art performance in motion planning on nuScenes as well as competitive results on the Waymo Open Motion Dataset (WOMD). EMMA also yields competitive results for camera-primary 3D object detection on the Waymo Open Dataset (WOD). We show that co-training EMMA with planner trajectories, object detection, and road graph tasks yields improvements across all three domains, highlighting EMMA's potential as a generalist model for autonomous driving applications. However, EMMA also exhibits certain limitations: it can process only a small amount of image frames, does not incorporate accurate 3D sensing modalities like LiDAR or radar and is computationally expensive. We hope that our results will inspire further research to mitigate these issues and to further evolve the state of the art in autonomous driving model architectures.

📄 PDF Abstract BibTeX arXiv:2410.23262

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous DrivingLarge Language ModelMotion Planningobject-detectionObject DetectionWorld Knowledge

Similar Papers 제목 키워드 기반

LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving

2025-05-01 · Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu

Vision-Language Models (VLMs) have demonstrated significant potential for end-to-end autonomous driving. However, fully exploiting their capabilities for safe and reliable vehicle control remains an open research challen…

Autonomous Driving

OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving

2024-12-19 · Shuo Xing, Chengyuan Qian, Yuping Wang, Hongyuan Hua 외

Since the advent of Multimodal Large Language Models (MLLMs), they have made a significant impact across a wide range of real-world applications, particularly in Autonomous Driving (AD). Their ability to process complex …

Autonomous Driving

A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving

2026-01-07 · Liangdong Zhang, Yiming Nie, Haoyang Li, Fanjie Kong 외 arxiv

Efficient trajectory planning in off-road terrains presents a formidable challenge for autonomous vehicles, often necessitating complex multi-step pipelines. However, traditional approaches exhibit limited adaptability i…

Semantic SegmentationAutonomous VehiclesTrajectory PlanningAutonomous Driving

XVTP3D: Cross-view Trajectory Prediction Using Shared 3D Queries for Autonomous Driving

2023-08-17 · Zijian Song, Huikun Bi, Ruisi Zhang, Tianlu Mao 외

Trajectory prediction with uncertainty is a critical and challenging task for autonomous driving. Nowadays, we can easily access sensor data represented in multiple views. However, cross-view consistency has not been eva…

Autonomous DrivingPredictionTrajectory Prediction

A Survey on Multimodal Large Language Models for Autonomous Driving

2023-11-21 · Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye 외

With the emergence of Large Language Models (LLMs) and Vision Foundation Models (VFMs), multimodal AI systems benefiting from large models have the potential to equally perceive the real world, make decisions, and contro…

Autonomous Driving