paper-with-me

홈 › Papers

MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases

2025-09-25 · Ziang Luo, Kangan Qian, Jiahua Wang, Yuechen Luo, Jinyu Miao, Zheng Fu, Yunlong Wang, Sicong Jiang, Zilin Huang, Yifei Hu, Yuhao Yang, Hao Ye, Mengmeng Yang, Xiaojian Dong, Kun Jiang, Diange Yang arxiv

Vision-Language Models(VLMs) have demonstrated significant potential for end-to-end autonomous driving, yet a substantial gap remains between their current capabilities and the reliability necessary for real-world deployment. A critical challenge is their fragility, characterized by hallucinations and poor generalization in out-of-distribution (OOD) scenarios. To bridge this gap, we introduce MTRDrive, a novel framework that integrates procedural driving experiences with a dynamic toolkit to enhance generalization and proactive decision-making. MTRDrive addresses these limitations through a closed-loop system that combines a memory-based experience retrieval mechanism with dynamic toolkits. This synergy enables the model to interact more effectively with its environment, improving both reasoning and decision-making capabilities with the help of our memory-tool synergistic reasoning. Additionally, we introduce a new benchmark based on complex Roadwork construction scenarios to rigorously evaluate zero-shot generalization. Extensive experiments demonstrate the superior effectiveness of our approach. On the public NAVSIM benchmark, our 3B-parameter MTRDrive model achieves an exceptional PDMS of 88.3 without chain-of-thought and sets a state-of-the-art performance bar on high-level planning, with a driving metric score of 79.8\% and a planning accuracy of 82.6\%. Rigorous zero-shot evaluation on the new Roadwork-VLM benchmark shows a strong ability to reason robustly in unseen scenarios, achieving a driving metric score of 80.2\%. These results highlight MTRDrive's potential to advance autonomous driving toward safer and more reliable systems.

📄 PDF Abstract BibTeX arXiv:2509.20843

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationAutonomous Driving

Similar Papers 제목 키워드 기반

Towards Efficient Agents: A Co-Design of Inference Architecture and System

2025-12-20 · Weizhe Lin, Hui-Ling Zhen, Shuai Yang, Xian Wang 외 arxiv

The rapid development of large language model (LLM)-based agents has unlocked new possibilities for autonomous multi-turn reasoning and tool-augmented decision-making. However, their real-world deployment is hindered by …

A collaborative agent with two lightweight synergistic models for autonomous crystal materials research

2026-04-13 · Tongyu Shi, Yutang Li, Zhanyuan Li, Qian Liu 외 arxiv

Current large language models require hundreds of billions of parameters yet struggle with domain-specific reasoning and tool coordination in materials science. Here, we present MatBrain, a lightweight collaborative agen…

TA-Mem: Tool-Augmented Autonomous Memory Retrieval for LLM in Long-Term Conversational QA

2026-03-10 · Mengwei Yuan, Jianan Liu, Jing Yang, Xianyou Li 외 arxiv

Large Language Model (LLM) has exhibited strong reasoning ability in text-based contexts across various domains, yet the limitation of context window poses challenges for the model on long-range inference tasks and neces…

Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning

2026-01-07 · Jinyang Wu, Guocheng Zhai, Ruihan Jin, Jiahao Yuan 외 arxiv

The integration of large language models (LLMs) with external tools has significantly expanded the capabilities of AI agents. However, as the diversity of both LLMs and tools increases, selecting the optimal model-tool c…

Visual Reasoning

KG-Agent: An Efficient Autonomous Agent Framework for Complex Reasoning over Knowledge Graph

2024-02-17 · Jinhao Jiang, Kun Zhou, Wayne Xin Zhao, Yang song 외

In this paper, we aim to improve the reasoning ability of large language models (LLMs) over knowledge graphs (KGs) to answer complex questions. Inspired by existing methods that design the interaction strategy between LL…

Knowledge Graphs