Explanation for Trajectory Planning using Multi-modal Large Language Model for Autonomous Driving
End-to-end style autonomous driving models have been developed recently. These models lack interpretability of decision-making process from perception to control of the ego vehicle, resulting in anxiety for passengers. To alleviate it, it is effective to build a model which outputs captions describing future behaviors of the ego vehicle and their reason. However, the existing approaches generate reasoning text that inadequately reflects the future plans of the ego vehicle, because they train models to output captions using momentary control signals as inputs. In this study, we propose a reasoning model that takes future planning trajectories of the ego vehicle as inputs to solve this limitation with the dataset newly collected.
Code (0)
등록된 구현이 없습니다.
Tasks
Autonomous DrivingDecision MakingLanguage ModelingLanguage ModellingLarge Language ModelTrajectory PlanningSimilar Papers 제목 키워드 기반
Generalized Trajectory Scoring for End-to-end Multimodal Planning
End-to-end multi-modal planning is a promising paradigm in autonomous driving, enabling decision-making with diverse trajectory candidates. A key component is a robust trajectory scorer capable of selecting the optimal t…
Autonomous DrivingDomain GeneralizationNavSimScene Induced Multi-Modal Trajectory Forecasting via Planning
We address multi-modal trajectory forecasting of agents in unknown scenes by formulating it as a planning problem. We present an approach consisting of three models; a goal prediction model to identify potential goals of…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Trajectory ForecastingCogDrive: Cognition-Driven Multimodal Prediction-Planning Fusion for Safe Autonomy
Safe autonomous driving in mixed traffic requires a unified understanding of multimodal interactions and dynamic planning under uncertainty. Existing learning based approaches struggle to capture rare but safety critical…
Trajectory PredictionAutonomous DrivingA Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a unified multimodal framework. However, …
Visual Question AnsweringTrajectory PlanningAutonomous DrivingDecision MakingMATS: An Interpretable Trajectory Forecasting Representation for Planning and Control
Reasoning about human motion is a core component of modern human-robot interactive systems. In particular, one of the main uses of behavior prediction in autonomous systems is to inform robot motion planning and control.…
Autonomous DrivingComputational EfficiencyMotion PlanningTrajectory Forecasting