paper-with-me

Papers

DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving

2025-11-25 · Haibo HU, Lianming Huang, Nan Guan, Chun Jason Xue arxiv

Vision-Language Action (VLA) models unify perception, reasoning, and trajectory generation for autonomous driving, but suffer from significant inference latency due to deep transformer stacks. We present DeeAD, a training-free, action-guided early-exit framework that accelerates VLA planning by evaluating the physical feasibility of intermediate trajectories. Instead of relying on confidence scores, DeeAD terminates inference when predicted trajectories align with lightweight planning priors (e.g., Navigation or Low-precision Planning) within a tolerable deviation (<2m). To improve efficiency, we introduce a multi-hop controller that adaptively skips redundant layers based on the change rate of scores. DeeAD integrates into existing VLA models, such as ORION, without requiring retraining. Experiments on the Bench2Drive benchmark demonstrate up to 28% transformer-layer sparsity and 29% latency reduction, while preserving planning quality and safety.

📄 PDF Abstract BibTeX arXiv:2511.20720

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language Model

2022-11-21 · CVPR 2023 1 · Shengkun Tang, Yaqing Wang, Zhenglun Kong, Tianchi Zhang 외

Large-scale Transformer models bring significant improvements for various downstream vision language tasks with a unified architecture. The performance improvements come with increasing model size, resulting in slow infe…

DecoderLanguage ModelingLanguage Modelling

LGViT: Dynamic Early Exiting for Accelerating Vision Transformer

2023-08-01 · Guanyu Xu, Jiawei Hao, Li Shen, Han Hu 외

Recently, the efficient deployment and acceleration of powerful vision transformers (ViTs) on resource-limited edge devices for providing multimedia services have become attractive tasks. Although early exiting is a feas…

DEED: Dynamic Early Exit on Decoder for Accelerating Encoder-Decoder Transformer Models

2023-11-15 · Peng Tang, Pengkai Zhu, Tian Li, Srikar Appalaraju 외

Encoder-decoder transformer models have achieved great success on various vision-language (VL) tasks, but they suffer from high inference latency. Typically, the decoder takes up most of the latency because of the auto-r…

Decoder

Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving

2025-10-02 · Haibo Hu, Lianming Huang, Xinyu Wang, Yufei Cui 외 arxiv

Vision-Language Models (VLMs) are increasingly applied in autonomous driving for unified perception and reasoning, but high inference latency hinders real-time deployment. Early-exit reduces latency by terminating infere…

Autonomous Driving

Single-Layer Vision Transformers for More Accurate Early Exits with Less Overhead

2021-05-19 · Arian Bakhtiarnia, Qi Zhang, Alexandros Iosifidis

Deploying deep learning models in time-critical applications with limited computational resources, for instance in edge computing systems and IoT networks, is a challenging task that often relies on dynamic inference met…

Audio ClassificationCrowd CountingEdge-computing