MTDrive: Multi-turn Interactive Reinforcement Learning for Autonomous Driving
Trajectory planning is a core task in autonomous driving, requiring the prediction of safe and comfortable paths across diverse scenarios. Integrating Multi-modal Large Language Models (MLLMs) with Reinforcement Learning (RL) has shown promise in addressing "long-tail" scenarios. However, existing methods are constrained to single-turn reasoning, limiting their ability to handle complex tasks requiring iterative refinement. To overcome this limitation, we present MTDrive, a multi-turn framework that enables MLLMs to iteratively refine trajectories based on environmental feedback. MTDrive introduces Multi-Turn Group Relative Policy Optimization (mtGRPO), which mitigates reward sparsity by computing relative advantages across turns. We further construct an interactive trajectory understanding dataset from closed-loop simulation to support multi-turn training. Experiments on the NAVSIM benchmark demonstrate superior performance compared to existing methods, validating the effectiveness of our multi-turn reasoning paradigm. Additionally, we implement system-level optimizations to reduce data transfer overhead caused by high-resolution images and multi-turn sequences, achieving 2.5x training throughput. Our data, models, and code will be made available soon.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningTrajectory PlanningAutonomous DrivingSimilar Papers 제목 키워드 기반
MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning
Medical image segmentation is evolving from task-specific models toward generalizable frameworks. Recent research leverages Multi-modal Large Language Models (MLLMs) as autonomous agents, employing reinforcement learning…
Medical Image SegmentationInteractive SegmentationReinforcement LearningMulti-task Safe Reinforcement Learning for Navigating Intersections in Dense Traffic
Multi-task intersection navigation including the unprotected turning left, turning right, and going straight in dense traffic is still a challenging task for autonomous driving. For the human driver, the negotiation skil…
Autonomous Drivingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
Effective interactive tool use requires agents to master Tool Integrated Reasoning (TIR): a complex process involving multi-turn planning and long-context dialogue management. To train agents for this dynamic process, pa…
Reinforcement LearningMathematical ReasoningDeep Interactive Reinforcement Learning for Path Following of Autonomous Underwater Vehicle
Autonomous underwater vehicle (AUV) plays an increasingly important role in ocean exploration. Existing AUVs are usually not fully autonomous and generally limited to pre-planning or pre-programming tasks. Reinforcement …
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving
Multi-agent interaction is a fundamental aspect of autonomous driving in the real world. Despite more than a decade of research and development, the problem of how to competently interact with diverse road users in diver…
Autonomous DrivingMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning (RL)