paper-with-me

홈 › Papers

Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

2025-05-29 · Haohan Chi, Huan-ang Gao, Ziming Liu, Jianing Liu, Chenyu Liu, Jinwei Li, Kaisen Yang, Yangcheng Yu, Zeda Wang, Wenyi Li, Leichen Wang, Xingtao Hu, Hao Sun, Hang Zhao, Hao Zhao

Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution is the Impromptu VLA Dataset: over 80,000 meticulously curated video clips, distilled from over 2M source clips sourced from 8 open-source large-scale datasets. This dataset is built upon our novel taxonomy of four challenging unstructured categories and features rich, planning-oriented question-answering annotations and action trajectories. Crucially, experiments demonstrate that VLAs trained with our dataset achieve substantial performance gains on established benchmarks--improving closed-loop NeuroNCAP scores and collision rates, and reaching near state-of-the-art L2 accuracy in open-loop nuScenes trajectory prediction. Furthermore, our Q&A suite serves as an effective diagnostic, revealing clear VLM improvements in perception, prediction, and planning. Our code, data and models are available at https://github.com/ahydchh/Impromptu-VLA.

📄 PDF Abstract BibTeX arXiv:2505.23757

Code (1)

ahydchh/impromptu-vla 공식 구현 pytorch

Tasks

Autonomous DrivingDiagnosticQuestion AnsweringTrajectory PredictionVision-Language-Action

Similar Papers 제목 키워드 기반

A Measurement of Social Capital in an Open Source Software Project

2019-11-22 · Saad Alqithami, Musaad Alzahrani, Fahad AlGhamdi, Rahmat Budiarto 외

The paper provides an understanding of social capital in organizations that are open membership multi-agent systems with an emphasis in our formulation on the dynamic network of social interaction that, in part, elucidat…

Impromptu Cybercrime Euphemism Detection

2024-12-02 · Xiang Li, Yucheng Zhou, Laiping Zhao, Jing Li 외

Detecting euphemisms is essential for content security on various social media platforms, but existing methods designed for detecting euphemisms are ineffective in impromptu euphemisms. In this work, we make a first atte…

An Inversion-Based Learning Approach for Improving Impromptu Trajectory Tracking of Robots with Non-Minimum Phase Dynamics

2017-09-13 · Siqi Zhou, Mohamed K. Helwa, Angela P. Schoellig

This paper presents a learning-based approach for impromptu trajectory tracking for non-minimum phase systems, i.e., systems with unstable inverse dynamics. Inversion-based feedforward approaches are commonly used for im…

VaViM and VaVAM: Autonomous Driving through Video Generative Modeling

2025-02-21 · Florent Bartoccioni, Elias Ramzi, Victor Besnier, Shashanka Venkataramanan 외

We explore the potential of large-scale generative video models for autonomous driving, introducing an open-source auto-regressive video model (VaViM) and its companion video-action model (VaVAM) to investigate how video…

Autonomous DrivingImitation Learning

SoTaNa: The Open-Source Software Development Assistant

2023-08-25 · Ensheng Shi, Fengji Zhang, Yanlin Wang, Bei Chen 외

Software development plays a crucial role in driving innovation and efficiency across modern societies. To meet the demands of this dynamic field, there is a growing need for an effective software development assistant. …

Code SummarizationGPUparameter-efficient fine-tuning