paper-with-me

홈 › Papers

Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving

2025-09-29 · Sheng Yang, Tong Zhan, Guancheng Chen, Yanfeng Lu, Jian Wang arxiv

In this work, we reconceptualize autonomous driving as a generalized language problem and formulate the trajectory planning task as next waypoint prediction. We introduce Max-V1, a novel framework for one-stage end-to-end autonomous driving, named in tribute to the renowned Dutch racing driver Max Verstappen. Our framework presents a single-pass generation paradigm that aligns with the inherent sequentiality of driving. This approach leverages the generative capacity of the Vision-Language Model (VLM) to enable end-to-end trajectory prediction directly from front-view camera input. The efficacy of this method is underpinned by a principled supervision strategy derived from statistical modeling. This provides a well-defined learning objective, which makes the framework highly amenable to mastering complex driving policies through imitation learning from large-scale expert demonstrations. Empirically, our method achieves state-of-the-art performance on the nuScenes dataset, delivering an overall improvement of over 30% compared to prior baselines. Furthermore, it exhibits superior generalization performance on cross-domain datasets acquired from diverse vehicles, demonstrating notable potential for cross-vehicle robustness and adaptability. With these empirical strengths, this work introduces a model that enables fundamental driving behaviors, laying the foundation for the development of more capable self-driving agents. Code will be available upon publication.

📄 PDF Abstract BibTeX arXiv:2510.00060

Code (0)

등록된 구현이 없습니다.

Tasks

Trajectory PredictionTrajectory PlanningAutonomous Driving

Similar Papers 제목 키워드 기반

PIP: Detecting Adversarial Examples in Large Vision-Language Models via Attention Patterns of Irrelevant Probe Questions

2024-09-08 · Yudong Zhang, Ruobing Xie, Jiansheng Chen, Xingwu Sun 외

Large Vision-Language Models (LVLMs) have demonstrated their powerful multimodal capabilities. However, they also face serious safety problems, as adversaries can induce robustness issues in LVLMs through the use of well…

INN: A Method Identifying Clean-annotated Samples via Consistency Effect in Deep Neural Networks

2021-06-29 · Dongha Kim, Yongchan Choi, Kunwoong Kim, Yongdai Kim

In many classification problems, collecting massive clean-annotated data is not easy, and thus a lot of researches have been done to handle data with noisy labels. Most recent state-of-art solutions for noisy label probl…

Memorization

VALD: Multi-Stage Vision Attack Detection for Efficient LVLM Defense

2026-02-23 · Nadav Kadvil, Malak Fares, Ayellet Tal arxiv

Large Vision-Language Models (LVLMs) can be vulnerable to adversarial images that subtly bias their outputs toward plausible yet incorrect responses. We introduce a general, efficient, and training-free defense that comb…

CLIPCleaner: Cleaning Noisy Labels with CLIP

2024-08-19 · Chen Feng, Georgios Tzimiropoulos, Ioannis Patras

Learning with Noisy labels (LNL) poses a significant challenge for the Machine Learning community. Some of the most widely used approaches that select as clean samples for which the model itself (the in-training model) h…

Learning with noisy labels

Improving Chemical Understanding of LLMs via SMILES Parsing

2025-05-22 · Yunhui Jang, Jaehyung Kim, Sungsoo Ahn

Large language models (LLMs) are increasingly recognized as powerful tools for scientific discovery, particularly in molecular science. A fundamental requirement for these models is the ability to accurately understand m…

Graph Matchingscientific discovery