paper-with-me

홈 › Papers

Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions

2025-04-03 · Peijie Yu, Yifan Yang, Jinjian Li, Zelong Zhang, Haorui Wang, Xiao Feng, Feng Zhang

Large language models (LLMs) demonstrate strong potential as agents for tool invocation due to their advanced comprehension and planning capabilities. Users increasingly rely on LLM-based agents to solve complex missions through iterative interactions. However, existing benchmarks predominantly access agents in single-mission scenarios, failing to capture real-world complexity. To bridge this gap, we propose the Multi-Mission Tool Bench. In the benchmark, each test case comprises multiple interrelated missions. This design requires agents to dynamically adapt to evolving demands. Moreover, the proposed benchmark explores all possible mission-switching patterns within a fixed mission number. Specifically, we propose a multi-agent data generation framework to construct the benchmark. We also propose a novel method to evaluate the accuracy and efficiency of agent decisions with dynamic decision trees. Experiments on diverse open-source and closed-source LLMs reveal critical factors influencing agent robustness and provide actionable insights to the tool invocation society.

📄 PDF Abstract BibTeX arXiv:2504.02623

Code (0)

등록된 구현이 없습니다.

Tasks

AI Agent

Similar Papers 제목 키워드 기반

Residual Error: a New Performance Measure for Adversarial Robustness

2021-06-18 · Hossein Aboutalebi, Mohammad Javad Shafiee, Michelle Karg, Christian Scharfenberger 외

Despite the significant advances in deep learning over the past decade, a major challenge that limits the wide-spread adoption of deep learning has been their fragility to adversarial attacks. This sensitivity to making …

Adversarial Robustnessimage-classificationImage Classification

Assessing the Robustness of Intelligence-Driven Reinforcement Learning

2023-11-15 · Lorenzo Nodari, Federico Cerutti

Robustness to noise is of utmost importance in reinforcement learning systems, particularly in military contexts where high stakes and uncertain environments prevail. Noise and uncertainty are inherent features of milita…

Decision Makingreinforcement-learningReinforcement Learning

$\text{R}^2$-Bench: Benchmarking the Robustness of Referring Perception Models under Perturbations

2024-03-07 · Xiang Li, Kai Qiu, Jinglu Wang, Xiaohao Xu 외

Referring perception, which aims at grounding visual objects with multimodal referring guidance, is essential for bridging the gap between humans, who provide instructions, and the environment where intelligent systems p…

Benchmarking

$α^3$-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks

2026-01-01 · Mohamed Amine Ferrag, Abderrahmane Lakas, Merouane Debbah arxiv

Large Language Models (LLMs) are increasingly used as high level controllers for autonomous Unmanned Aerial Vehicle (UAV) missions. However, existing evaluations rarely assess whether such agents remain safe, protocol co…

Assessing Visually-Continuous Corruption Robustness of Neural Networks Relative to Human Performance

2024-02-29 · Huakun Shen, Boyue Caroline Hu, Krzysztof Czarnecki, Lina Marsso 외

While Neural Networks (NNs) have surpassed human accuracy in image classification on ImageNet, they often lack robustness against image corruption, i.e., corruption robustness. Yet such robustness is seemingly effortless…

Data Augmentationimage-classificationImage Classification