ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
Large language models (LLMs) have demonstrated strong capabilities in using external tools to address user inquiries. However, most existing evaluations assume tool use in short contexts, offering limited insight into model behavior during realistic long-term interactions. To fill this gap, we introduce ToolHaystack, a benchmark for testing the tool use capabilities in long-term interactions. Each test instance in ToolHaystack includes multiple tasks execution contexts and realistic noise within a continuous conversation, enabling assessment of how well models maintain context and handle various disruptions. By applying this benchmark to 14 state-of-the-art LLMs, we find that while current models perform well in standard multi-turn settings, they often significantly struggle in ToolHaystack, highlighting critical gaps in their long-term robustness not revealed by previous tool benchmarks.
Code (1)
Similar Papers 제목 키워드 기반
Adaptive Stress Testing for Adversarial Learning in a Financial Environment
We demonstrate the use of Adaptive Stress Testing to detect and address potential vulnerabilities in a financial environment. We develop a simplified model for credit card fraud detection that utilizes a linear regressio…
Fraud Detectionregressionreinforcement-learningReinforcement Learning (RL)Liquidity Stress Testing in Asset Management -- Part 3. Managing the Asset-Liability Liquidity Risk
This article is part of a comprehensive research project on liquidity risk in asset management, which can be divided into three dimensions. The first dimension covers the modeling of the liability liquidity risk (or fund…
Asset ManagementManagementIntenTest: Stress Testing for Intent Integrity in API-Calling LLM Agents
LLM agents are increasingly deployed to automate real-world tasks by invoking APIs through natural language instructions. While powerful, they often suffer from misinterpretation of user intent, leading to the agent's ac…
software testingHacking, The Lazy Way: LLM Augmented Pentesting
In our research, we introduce a new concept called "LLM Augmented Pentesting" demonstrated with a tool named "Pentest Copilot," that revolutionizes the field of ethical hacking by integrating Large Language Models (LLMs)…
RAGRetrieval-augmented GenerationCausal Data Science for Financial Stress Testing
The most recent financial upheavals have cast doubt on the adequacy of some of the conventional quantitative risk management strategies, such as VaR (Value at Risk), in many common situations. Consequently, there has bee…
Management