paper-with-me

홈 › Papers

DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment

2025-10-20 · Yu Gao, Anqing Jiang, Yiru Wang, Wang Jijun, Hao Jiang, Zhigang Sun, Heng Yuwen, Wang Shuo, Hao Zhao, Sun Hao arxiv

Conventional end-to-end (E2E) driving models are effective at generating physically plausible trajectories, but often fail to generalize to long-tail scenarios due to the lack of essential world knowledge to understand and reason about surrounding environments. In contrast, Vision-Language-Action (VLA) models leverage world knowledge to handle challenging cases, but their limited 3D reasoning capability can lead to physically infeasible actions. In this work we introduce DiffVLA++, an enhanced autonomous driving framework that explicitly bridges cognitive reasoning and E2E planning through metric-guided alignment. First, we build a VLA module directly generating semantically grounded driving trajectories. Second, we design an E2E module with a dense trajectory vocabulary that ensures physical feasibility. Third, and most critically, we introduce a metric-guided trajectory scorer that guides and aligns the outputs of the VLA and E2E modules, thereby integrating their complementary strengths. The experiment on the ICCV 2025 Autonomous Grand Challenge leaderboard shows that DiffVLA++ achieves EPDMS of 49.12.

📄 PDF Abstract BibTeX arXiv:2510.17148

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving

2025-05-26 · Anqing Jiang, Yu Gao, Zhigang Sun, Yiru Wang 외

Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ulti…

Autonomous DrivingDiversityLanguage ModelingLanguage Modelling

Bridging Social Psychology and LLM Reasoning: Conflict-Aware Meta-Review Generation via Cognitive Alignment

2025-03-18 · Wei Chen, Han Ding, Meng Yuan, Zhao Zhang 외

The rapid growth of scholarly submissions has overwhelmed traditional peer review systems, driving the need for intelligent automation to preserve scientific rigor. While large language models (LLMs) show promise in auto…

Review Generation

CHARMS: A Cognitive Hierarchical Agent for Reasoning and Motion Stylization in Autonomous Driving

2025-04-03 · Jingyi Wang, DuanFeng Chu, Zejian Deng, LiPing Lu 외

To address the challenge of insufficient interactivity and behavioral diversity in autonomous driving decision-making, this paper proposes a Cognitive Hierarchical Agent for Reasoning and Motion Stylization (CHARMS). By …

Autonomous DrivingDecision MakingDeep Reinforcement LearningDiversity

Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning

2025-02-19 · Rui Zhao, Qirui Yuan, Jinyu Li, Haofeng Hu 외

End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an important part of Embodied AI. Despite successes in applying Multimodal Large Language Models (MLLMs) for high-lev…

Autonomous DrivingBench2DriveMotion PlanningQuestion Answering+3

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models

2026-05-24 · Boyu Li, Chaoyi Xu, Haoqi Yuan, Xinrun Xu 외 arxiv

Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-trained on large and diverse datasets, they typically rely on embodim…