paper-with-me

홈 › Papers

Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models

2025-10-20 · Katie Luo, Jingwei Ji, Tong He, Runsheng Xu, Yichen Xie, Dragomir Anguelov, Mingxing Tan arxiv

Current autonomous driving systems rely on specialized models for perceiving and predicting motion, which demonstrate reliable performance in standard conditions. However, generalizing cost-effectively to diverse real-world scenarios remains a significant challenge. To address this, we propose Plug-and-Forecast (PnF), a plug-and-play approach that augments existing motion forecasting models with multimodal large language models (MLLMs). PnF builds on the insight that natural language provides a more effective way to describe and handle complex scenarios, enabling quick adaptation to targeted behaviors. We design prompts to extract structured scene understanding from MLLMs and distill this information into learnable embeddings to augment existing behavior prediction models. Our method leverages the zero-shot reasoning capabilities of MLLMs to achieve significant improvements in motion prediction performance, while requiring no fine-tuning -- making it practical to adopt. We validate our approach on two state-of-the-art motion forecasting models using the Waymo Open Motion Dataset and the nuScenes Dataset, demonstrating consistent performance improvements across both benchmarks.

📄 PDF Abstract BibTeX arXiv:2510.17274

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingAutonomous DrivingMotion Forecasting

Similar Papers 제목 키워드 기반

Leveraging Self-Paced Curriculum Learning for Enhanced Modality Balance in Multimodal Conversational Emotion Recognition

2026-05-20 · Phuong-Anh Nguyen, The-Son Le, Duc-Trong Le, Cam-Van Thi Nguyen arxiv

Multimodal Emotion Recognition in Conversations (MERC) is a crucial task for understanding human interactions, where multimodal approaches integrating language, facial expressions, and vocal tone have achieved significan…

Multimodal Emotion Recognition

MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls

2024-07-30 · Yuxuan Bian, Ailing Zeng, Xuan Ju, Xian Liu 외

Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generatio…

Gesture GenerationMotion GenerationMotion Synthesismultimodal generation

Context Matters: Leveraging Contextual Features for Time Series Forecasting

2024-10-16 · Sameep Chattopadhyay, Pulkit Paliwal, Sai Shankar Narasimhan, Shubhankar Agarwal 외

Time series forecasts are often influenced by exogenous contextual features in addition to their corresponding history. For example, in financial settings, it is hard to accurately predict a stock price without consideri…

ArticlesTime SeriesTime Series Forecasting

GS-FUSE: Granger-Supervised Gated Fusion and Multi-Granularity Alignment for Event-Driven Financial Forecasting

2026-05-27 · Yang Zhang, En Chun, Ziyun Mao, Yulu Wu 외 arxiv

Accurately forecasting the impact of salient financial events on markets is critical for investors and policymakers. However, existing multimodal time-series models typically fuse text and prices symmetrically, without a…

TaPD: Temporal-adaptive Progressive Distillation for Observation-Adaptive Trajectory Forecasting in Autonomous Driving

2026-03-06 · Mingyu Fan, Yi Liu, Hao Zhou, Deheng Qian 외 arxiv

Trajectory prediction is essential for autonomous driving, enabling vehicles to anticipate the motion of surrounding agents to support safe planning. However, most existing predictors assume fixed-length histories and su…

Trajectory ForecastingKnowledge DistillationTrajectory PredictionAutonomous Driving