paper-with-me

홈 › Papers

StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs

2025-03-26 · Zhicheng Guo, Sijie Cheng, Yuchen Niu, Hao Wang, Sicheng Zhou, Wenbing Huang, Yang Liu

The rapid advancement of large language models (LLMs) has spurred significant interest in tool learning, where LLMs are augmented with external tools to tackle complex tasks. However, existing tool environments face challenges in balancing stability, scalability, and realness, particularly for benchmarking purposes. To address this problem, we propose MirrorAPI, a novel framework that trains specialized LLMs to accurately simulate real API responses, effectively acting as "mirrors" to tool environments. Using a comprehensive dataset of request-response pairs from 7,000+ APIs, we employ supervised fine-tuning and chain-of-thought reasoning to enhance simulation fidelity. MirrorAPI achieves superior accuracy and stability compared to state-of-the-art methods, as demonstrated by its performance on the newly constructed MirrorAPI-Bench and its integration into StableToolBench.

📄 PDF Abstract BibTeX arXiv:2503.20527

Code (1)

thunlp-mt/stabletoolbench 공식 구현

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

2024-03-12 · Zhicheng Guo, Sijie Cheng, Hao Wang, Shihao Liang 외

Large Language Models (LLMs) have witnessed remarkable advancements in recent years, prompting the exploration of tool learning, which integrates LLMs with external tools to address diverse real-world challenges. Assessi…

Benchmarking

Self-Improving World Modelling with Latent Actions

2026-02-05 · Yifu Qiu, Zheng Zhao, Waylon Li, Yftah Ziser 외 arxiv

Internal modelling of the world -- predicting transitions between previous states $X$ and next states $Y$ under actions $Z$ -- is essential to reasoning and planning for LLMs and VLMs. Learning such models typically requ…

Reinforcement Learning

Budget-Constrained Agentic Large Language Models: Intention-Based Planning for Costly Tool Use

2026-02-12 · Hanbing Liu, Chunhao Tian, Nan An, Ziyuan Wang 외 arxiv

We study budget-constrained tool-augmented agents, where a large language model must solve multi-step tasks by invoking external tools under a strict monetary budget. We formalize this setting as sequential decision maki…

Decision Making

Reducing Tool Hallucination via Reliability Alignment

2024-12-05 · Hongshen Xu, Su Zhu, Zihan Wang, Hang Zheng 외

Large Language Models (LLMs) have extended their capabilities beyond language generation to interact with external systems through tool calling, offering powerful potential for real-world applications. However, the pheno…

HallucinationText Generation

On Using Curved Mirrors to Decrease Shadowing in VLC

2024-09-05 · Borja Genoves Guzman, Ana Garcia Armada, Maïté Brandt-Pearce

Visible light communication (VLC) complements radio frequency in indoor environments with large wireless data traffic. However, VLC is hindered by dramatic path losses when an opaque object is interposed between the tran…