paper-with-me

홈 › Papers

ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions

2025-05-29 · Beong-woo Kwak, Minju Kim, Dongha Lim, Hyungjoo Chae, Dongjin Kang, Sunghwan Kim, Dongil Yang, Jinyoung Yeo

Large language models (LLMs) have demonstrated strong capabilities in using external tools to address user inquiries. However, most existing evaluations assume tool use in short contexts, offering limited insight into model behavior during realistic long-term interactions. To fill this gap, we introduce ToolHaystack, a benchmark for testing the tool use capabilities in long-term interactions. Each test instance in ToolHaystack includes multiple tasks execution contexts and realistic noise within a continuous conversation, enabling assessment of how well models maintain context and handle various disruptions. By applying this benchmark to 14 state-of-the-art LLMs, we find that while current models perform well in standard multi-turn settings, they often significantly struggle in ToolHaystack, highlighting critical gaps in their long-term robustness not revealed by previous tool benchmarks.

📄 PDF Abstract BibTeX arXiv:2505.23662

Code (1)

bwookwak/toolhaystack 공식 구현

Similar Papers 제목 키워드 기반

Adaptive Stress Testing for Adversarial Learning in a Financial Environment

2021-07-08 · Khalid El-Awady

We demonstrate the use of Adaptive Stress Testing to detect and address potential vulnerabilities in a financial environment. We develop a simplified model for credit card fraud detection that utilizes a linear regressio…

Fraud Detectionregressionreinforcement-learningReinforcement Learning (RL)

Liquidity Stress Testing in Asset Management -- Part 3. Managing the Asset-Liability Liquidity Risk

2021-10-04 · Thierry Roncalli

This article is part of a comprehensive research project on liquidity risk in asset management, which can be divided into three dimensions. The first dimension covers the modeling of the liability liquidity risk (or fund…

Asset ManagementManagement

IntenTest: Stress Testing for Intent Integrity in API-Calling LLM Agents

2025-06-09 · Shiwei Feng, Xiangzhe Xu, Xuan Chen, Kaiyuan Zhang 외

LLM agents are increasingly deployed to automate real-world tasks by invoking APIs through natural language instructions. While powerful, they often suffer from misinterpretation of user intent, leading to the agent's ac…

software testing

Hacking, The Lazy Way: LLM Augmented Pentesting

2024-09-14 · Dhruva Goyal, Sitaraman Subramanian, Aditya Peela, Nisha P. Shetty

In our research, we introduce a new concept called "LLM Augmented Pentesting" demonstrated with a tool named "Pentest Copilot," that revolutionizes the field of ethical hacking by integrating Large Language Models (LLMs)…

RAGRetrieval-augmented Generation

Causal Data Science for Financial Stress Testing

2017-03-08 · Gelin Gao, Bud Mishra, Daniele Ramazzotti

The most recent financial upheavals have cast doubt on the adequacy of some of the conventional quantitative risk management strategies, such as VaR (Value at Risk), in many common situations. Consequently, there has bee…

Management