paper-with-me

Papers

LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

2026-05-21 · Xiaodong Mei, Diankun Zhang, Hongwei Xie, Guang Chen, Hangjun Ye, Dan Xu arxiv

Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene understanding and reasoning capabilities. Recent attempts to incorporate dense visual supervision via world modeling often overemphasize pixel-level image reconstruction, neglecting semantically meaningful scene representation learning. In this work, we propose LVDrive, a Latent Visual representation enhanced VLA framework for autonomous driving. LVDrive introduces a future scene prediction task into the VLA paradigm, where future representations are learned entirely in a high-level latent space under auxiliary supervision from a pretrained vision backbone. Departing from inefficient autoregressive generation, we jointly model future scene and motion prediction within a unified embedding space, processed in a single forward pass to conduct the future-aware reasoning. We further design a two-stage trajectory decoding strategy that explicitly leverages the learned latent future representations to refine trajectory generation. Extensive experiments on the challenging Bench2Drive benchmark demonstrate that LVDrive achieves significant improvements in closed-loop driving performance, outperforming both action supervised methods and image-reconstruction-based world model approaches.

📄 PDF Abstract BibTeX arXiv:2605.22089

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningImage ReconstructionScene UnderstandingAutonomous Driving

Similar Papers 제목 키워드 기반

Vision-Enhanced Time Series Forecasting via Latent Diffusion Models

2025-02-16 · Weilin Ruan, Siru Zhong, Haomin Wen, Yuxuan Liang

Diffusion models have recently emerged as powerful frameworks for generating high-quality images. While recent studies have explored their application to time series forecasting, these approaches face significant challen…

Image ReconstructionTime SeriesTime Series Forecasting

Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving

2024-05-02 · Zhenjiang Mao, Dong-You Jhong, Ao Wang, Ivan Ruchkin

Out-of-distribution (OOD) detection is essential in autonomous driving, to determine when learning-based components encounter unexpected inputs. Traditional detectors typically use encoder models with fixed settings, thu…

Anomaly DetectionAutonomous DrivingOut-of-Distribution DetectionOut of Distribution (OOD) Detection

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs

2026-05-27 · Qiaoru Li, Shaotian Liang, Jintao Chen, Haoran Sun 외 arxiv

Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead of chain-of-thought for medical VQA. However, existing methods suffer …

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs

2026-01-31 · Benno Krojer, Shravan Nayak, Oscar Mañas, Vaibhav Adlakha 외 arxiv

Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM. Intriguingly, this mapping can be as sim…

Combined CNN Transformer Encoder for Enhanced Fine-grained Human Action Recognition

2022-08-03 · Mei Chee Leong, Haosong Zhang, Hui Li Tan, Liyuan Li 외

Fine-grained action recognition is a challenging task in computer vision. As fine-grained datasets have small inter-class variations in spatial and temporal space, fine-grained action recognition model requires good temp…

Action RecognitionAttributeFine-grained Action RecognitionTemporal Action Localization