paper-with-me

홈 › Papers

Tools Fail: Detecting Silent Errors in Faulty Tools

2024-06-27 · Jimin Sun, So Yeon Min, Yingshan Chang, Yonatan Bisk

Tools have become a mainstay of LLMs, allowing them to retrieve knowledge not in their weights, to perform tasks on the web, and even to control robots. However, most ontologies and surveys of tool-use have assumed the core challenge for LLMs is choosing the tool. Instead, we introduce a framework for tools more broadly which guides us to explore a model's ability to detect "silent" tool errors, and reflect on how to plan. This more directly aligns with the increasingly popular use of models as tools. We provide an initial approach to failure recovery with promising results both on a controlled calculator setting and embodied agent planning.

📄 PDF Abstract BibTeX arXiv:2406.19228

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Thales: Formulating and Estimating Architectural Vulnerability Factors for DNN Accelerators

2022-12-05 · Abhishek Tyagi, Yiming Gan, Shaoshan Liu, Bo Yu 외

As Deep Neural Networks (DNNs) are increasingly deployed in safety critical and privacy sensitive applications such as autonomous driving and biometric authentication, it is critical to understand the fault-tolerance nat…

Autonomous Driving

On the Use of CVRP to Diagnose Faulty Elements in Antenna Arrays

2025-05-13 · Alejandro Antón Ruiz, John Kvarnstrand, Klas Arvidsson, Andrés Alayón Glazunov

This paper investigates the application of Constrained-View Radiated Power (CVRP) for diagnosing phased array element failures, specifically focusing on on-off element failure. CVRP, similar to Partial Radiated Power (PR…

Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework

2025-08-29 · Nils Dycke, Iryna Gurevych arxiv

Large Language Models (LLMs) have great potential to accelerate and support scholarly peer review and are increasingly used as fully automatic review generators (ARGs). However, potential biases and systematic errors may…

Understanding Silent Failures in Medical Image Classification

2023-07-27 · Till J. Bungert, Levin Kobelke, Paul F. Jaeger

To ensure the reliable use of classification systems in medical applications, it is crucial to prevent silent failures. This can be achieved by either designing classifiers that are robust enough to avoid failures in the…

Classificationimage-classificationImage ClassificationMedical Image Classification

When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime

2026-06-12 · Wei Wu arxiv

LLM agent systems increasingly run as long-lived autonomous runtimes: scheduling jobs, calling tools, maintaining memory, and pushing results to humans. We present a longitudinal study of silent failures in one such syst…