paper-with-me

홈 › Papers

ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

2023-07-31 · Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, Maosong Sun

Despite the advancements of open-source large language models (LLMs), e.g., LLaMA, they remain significantly limited in tool-use capabilities, i.e., using external tools (APIs) to fulfill human instructions. The reason is that current instruction tuning largely focuses on basic language tasks but ignores the tool-use domain. This is in contrast to the excellent tool-use capabilities of state-of-the-art (SOTA) closed-source LLMs, e.g., ChatGPT. To bridge this gap, we introduce ToolLLM, a general tool-use framework encompassing data construction, model training, and evaluation. We first present ToolBench, an instruction-tuning dataset for tool use, which is constructed automatically using ChatGPT. Specifically, the construction can be divided into three stages: (i) API collection: we collect 16,464 real-world RESTful APIs spanning 49 categories from RapidAPI Hub; (ii) instruction generation: we prompt ChatGPT to generate diverse instructions involving these APIs, covering both single-tool and multi-tool scenarios; (iii) solution path annotation: we use ChatGPT to search for a valid solution path (chain of API calls) for each instruction. To enhance the reasoning capabilities of LLMs, we develop a novel depth-first search-based decision tree algorithm. It enables LLMs to evaluate multiple reasoning traces and expand the search space. Moreover, to evaluate the tool-use capabilities of LLMs, we develop an automatic evaluator: ToolEval. Based on ToolBench, we fine-tune LLaMA to obtain an LLM ToolLLaMA, and equip it with a neural API retriever to recommend appropriate APIs for each instruction. Experiments show that ToolLLaMA demonstrates a remarkable ability to execute complex instructions and generalize to unseen APIs, and exhibits comparable performance to ChatGPT. Our ToolLLaMA also demonstrates strong zero-shot generalization ability in an out-of-distribution tool-use dataset: APIBench.

📄 PDF Abstract BibTeX arXiv:2307.16789

Code (2)

openbmb/toolbench 공식 구현 pytorch
eachsheep/shortcutsbench

Tasks

Trajectory PlanningZero-shot Generalization

Similar Papers 제목 키워드 기반

AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls

2024-02-06 · Yu Du, Fangyun Wei, Hongyang Zhang

We introduce AnyTool, a large language model agent designed to revolutionize the utilization of a vast array of tools in addressing user queries. We utilize over 16,000 APIs from Rapid API, operating under the assumption…

Language ModelingLanguage ModellingLarge Language Model

Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning

2025-06-05 · Zhiyuan Ma, Jiayu Liu, Xianzhen Luo, Zhenya Huang 외

Empowering large language models (LLMs) with effective tool utilization capabilities is crucial for enabling AI agents to solve complex problems. However, current models face two major limitations: (1) unreliable tool pl…

Imitation Learning

Stepwise Self-Consistent Mathematical Reasoning with Large Language Models

2024-02-24 · Zilong Zhao, Yao Rong, Dongyang Guo, Emek Gözlüklü 외

Using Large Language Models for complex mathematical reasoning is difficult, primarily due to the complexity of multi-step reasoning. The main challenges of this process include (1) selecting critical intermediate result…

MathMathematical Reasoning

Exploring Autonomous Agents through the Lens of Large Language Models: A Review

2024-04-05 · Saikat Barua

Large Language Models (LLMs) are transforming artificial intelligence, enabling autonomous agents to perform diverse tasks across various domains. These agents, proficient in human-like text comprehension and generation,…

In-Context LearningReading Comprehension

MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools

2025-10-28 · Wenhao Wang, Peizhi Niu, Zhao Xu, Zhaoyu Chen 외 arxiv

Large Language Models (LLMs) increasingly rely on external tools to perform complex, realistic tasks, yet their ability to utilize the rapidly expanding Model Contextual Protocol (MCP) ecosystem remains limited. Existing…