paper-with-me

홈 › Papers

MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use

2025-08-22 · Fei Lei, Yibo Yang, Wenxiu Sun, Dahua Lin arxiv

Large Language Models (LLMs) are evolving from text generators into reasoning agents. This transition makes their ability to use external tools a critical capability. However, evaluating this skill presents a significant challenge. Existing benchmarks are often limited by their reliance on synthetic tools and severely constrained action spaces. To address these limitations, we introduce MCPVerse, an expansive, real-world benchmark for evaluating agentic tool use. MCPVerse integrates more than 550 real-world, executable tools to create an unprecedented action space exceeding 140k tokens, and employs outcome-based evaluation with real-time ground truth for time-sensitive tasks. We benchmarked the state-of-the-art LLMs across three modes (Oracle, Standard, and Max-Scale), revealing that while most models suffer performance degradation when confronted with larger tool sets, the agentic models, such as Claude-4-Sonnet, can effectively leverage expanded exploration spaces to improve accuracy. This finding not only exposes the limitations of state-of-the-art models in complex, real-world scenarios but also establishes MCPVerse as a critical benchmark for measuring and advancing agentic tool use capabilities.

📄 PDF Abstract BibTeX arXiv:2508.16260

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WorldAgents: Can Foundation Image Models be Agents for 3D World Models?

2026-03-20 · Ziya Erkoç, Angela Dai, Matthias Nießner arxiv

Given the remarkable ability of 2D foundation image models to generate high-fidelity outputs, we investigate a fundamental question: do 2D foundation image models inherently possess 3D world model capabilities? To answer…

3D ReconstructionImage Generation

Open-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents

2026-06-14 · Wasi Uddin Ahmad, Nikolai Ludwig, Somshubra Majumdar, Boris Ginsburg arxiv

The path toward autonomous software engineering is currently bottlenecked by a severe deficit of diverse, large-scale trajectory data. We address this by introducing \ourdataset, an expansive dataset of 207,489 agentic t…

HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization

2025-06-09 · Hongzheng Chen, Yingheng Wang, Yaohui Cai, Hins Hu 외

While Large Language Models (LLMs) have demonstrated significant advancements in reasoning and agent-based problem-solving, current evaluation methodologies fail to adequately assess their capabilities: existing benchmar…

Combinatorial OptimizationMemorization

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

2025-05-26 · Atsunori Moteki, Shoichi Masui, Fan Yang, Yueqi Song 외

This paper proposes FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agentic AI, they are required to monitor and report safety and health incidents, as w…

Universal Deep Research: Bring Your Own Model and Strategy

2025-08-29 · Peter Belcak, Pavlo Molchanov arxiv

Deep research tools are among the most impactful and most commonly encountered agentic systems today. We observe, however, that each deep research agent introduced so far is hard-coded to carry out a particular research …