paper-with-me

홈 › Papers

TheMCPCompany: Creating General-purpose Agents with Task-specific Tools

2025-10-22 · Reza Esfandiarpoor, Vishwas Suryanarayanan, Stephen H. Bach, Vishal Chowdhary, Anthony Aue arxiv

Since the introduction of the Model Context Protocol (MCP), the number of available tools for Large Language Models (LLMs) has increased significantly. These task-specific tool sets offer an alternative to general-purpose tools such as web browsers, while being easier to develop and maintain than GUIs. However, current general-purpose agents predominantly rely on web browsers for interacting with the environment. Here, we introduce TheMCPCompany, a benchmark for evaluating tool-calling agents on tasks that involve interacting with various real-world services. We use the REST APIs of these services to create MCP servers, which include over 18,000 tools. We also provide manually annotated ground-truth tools for each task. In our experiments, we use the ground truth tools to show the potential of tool-calling agents for both improving performance and reducing costs assuming perfect tool retrieval. Next, we explore agent performance using tool retrieval to study the real-world practicality of tool-based agents. While all models with tool retrieval perform similarly or better than browser-based agents, smaller models cannot take full advantage of the available tools through retrieval. On the other hand, GPT-5's performance with tool retrieval is very close to its performance with ground-truth tools. Overall, our work shows that the most advanced reasoning models are effective at discovering tools in simpler environments, but seriously struggle with navigating complex enterprise environments. TheMCPCompany reveals that navigating tens of thousands of tools and combining them in non-trivial ways to solve complex problems is still a challenging task for current models and requires both better reasoning and better retrieval models.

📄 PDF Abstract BibTeX arXiv:2510.19286

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning with Digital Agents: An Analysis based on the Activity Theory

2024-08-08 · Mateusz Dolata, Dzmitry Katsiuba, Natalie Wellnhammer, Gerhard Schwabe

Digital agents are considered a general-purpose technology. They spread quickly in private and organizational contexts, including education. Yet, research lacks a conceptual framing to describe interaction with such agen…

TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning

2025-05-26 · Yuhui Chen, Haoran Li, Zhennan Jiang, Haowei Wen 외

Developing scalable and generalizable reward engineering for reinforcement learning (RL) is crucial for creating general-purpose agents, especially in the challenging domain of robotic manipulation. While recent advances…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

2023-11-09 · Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang 외

LLaVA-Plus is a general-purpose multimodal assistant that expands the capabilities of large multimodal models. It maintains a skill repository of pre-trained vision and vision-language models and can activate relevant to…

Instruction FollowingLLM real-life tasksLMM real-life tasksRetrieval+1

Toward equipping Artificial Moral Agents with multiple ethical theories

2020-03-02 · George Rautenbach, C. Maria Keet

Artificial Moral Agents (AMA's) is a field in computer science with the purpose of creating autonomous machines that can make moral decisions akin to how humans do. Researchers have proposed theoretical means of creating…

Scaling Agents via Continual Pre-training

2025-09-16 · Liangcai Su, Zhen Zhang, Guangyu Li, Zhuo Chen 외 arxiv

Large language models (LLMs) have evolved into agentic systems capable of autonomous tool use and multi-step reasoning for complex problem-solving. However, post-training approaches building upon general-purpose foundati…