paper-with-me

Papers

Safe and Interpretable Multimodal Path Planning for Multi-Agent Cooperation

2026-02-22 · Haojun Shi, Suyu Ye, Katherine M. Guerrerio, Jianzhi Shen, Yifan Yin, Daniel Khashabi, Chien-Ming Huang, Tianmin Shu arxiv

Successful cooperation among decentralized agents requires each agent to quickly adapt its plan to the behavior of other agents. In scenarios where agents cannot confidently predict one another's intentions and plans, language communication can be crucial for ensuring safety. In this work, we focus on path-level cooperation in which agents must adapt their paths to one another in order to avoid collisions or perform physical collaboration such as joint carrying. In particular, we propose a safe and interpretable multimodal path planning method, CaPE (Code as Path Editor), which generates and updates path plans for an agent based on the environment and language communication from other agents. CaPE leverages a vision-language model (VLM) to synthesize a path editing program verified by a model-based planner, grounding communication to path plan updates in a safe and interpretable way. We evaluate our approach in diverse simulated and real-world scenarios, including multi-robot and human-robot cooperation in autonomous driving, household, and joint carrying tasks. Experimental results demonstrate that CaPE can be integrated into different robotic systems as a plug-and-play module, greatly enhancing a robot's ability to align its plan to language communication from other robots or humans. We also show that the combination of the VLM-based path editing program synthesis and model-based planning safety enables robots to achieve open-ended cooperation while maintaining safety and interpretability.

📄 PDF Abstract BibTeX arXiv:2602.19304

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingProgram Synthesis

Similar Papers 제목 키워드 기반

CogDrive: Cognition-Driven Multimodal Prediction-Planning Fusion for Safe Autonomy

2025-12-02 · Heye Huang, Yibin Yang, Mingfeng Fan, Haoran Wang 외 arxiv

Safe autonomous driving in mixed traffic requires a unified understanding of multimodal interactions and dynamic planning under uncertainty. Existing learning based approaches struggle to capture rare but safety critical…

Trajectory PredictionAutonomous Driving

When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models

2025-09-15 · Wei Cai, Shujuan Liu, Jian Zhao, Ziyan Shi 외 arxiv

Multimodal Large Language Models (MLLMs) are susceptible to the implicit reasoning risk, wherein innocuous unimodal inputs synergistically assemble into risky multimodal data that produce harmful outputs. We attribute th…

From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models

2025-06-02 · Yihong Tang, Ao Qu, Xujing Yu, Weipeng Deng 외

Urban and transportation research has long sought to uncover statistically meaningful relationships between key variables and societal outcomes such as road safety, to generate actionable insights that guide the planning…

Large Language ModelMultimodal Large Language Modelscientific discovery

Multimodal Trajectory Planning for Surface Vehicles using Turning Circle-based Control Barrier Functions

2026-08-20 · Changyu Lee arxiv

This paper presents a guide path-free multimodal trajectory planning framework for autonomous surface vehicles operating in dynamic environments. The proposed method integrates model predictive control (MPC) with a turni…

Computational EfficiencyTrajectory Planning

PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray Reasoning

2025-08-14 · Yushi Feng, Junye Du, Yingying Hong, Qifan Wang 외 arxiv

Existing tool-augmented agentic systems are limited in the real world by (i) black-box reasoning steps that undermine trust of decision-making and pose safety risks, (ii) poor multimodal integration, which is inherently …

Reinforcement LearningSemantic Similarity