paper-with-me

홈 › Papers

Invisible failures in human-AI interactions

2026-03-16 · Christopher Potts, Moritz Sudhof arxiv

AI systems fail silently far more often than they fail visibly. In an analysis of 100K human-AI interactions from the WildChat dataset, we find that 79% of AI failures are invisible: something went wrong but the user gave no overt indication that there was a problem. These invisible failures cluster into eight archetypes that help us characterize where and how AI systems are failing to meet users' needs. In addition, the archetypes show systematic co-occurrence patterns indicating higher-level failure types. To address the question of whether these archetypes will remain relevant as AI systems become more capable, we also created and annotated a counterfactual dataset in which WildChat's 2024-era responses are replaced by those from three present-day frontier LMs. This analysis indicates that failure rates have dropped substantially, but that the vast majority of failures remain invisible in our sense, and the distribution of failure archetypes seems stable. Finally, we illustrate how the archetypes help us to identify systematic and variable AI limitations across different usage domains. Overall, we argue that our invisible failure taxonomy can be a key component in reliable failure monitoring for product developers, scientists, and policy makers. Our code and data are available at https://github.com/bigspinai/bigspin-invisible-failure-archetypes

📄 PDF Abstract BibTeX arXiv:2603.15423

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild

2021-04-08 · ICCV 2021 10 · Vida Adeli, Mahsa Ehsanpour, Ian Reid, Juan Carlos Niebles 외

Joint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems. Predicting body dynamics requires capt…

Autonomous DrivingHuman-Object Interaction Detection

Ins-HOI: Instance Aware Human-Object Interactions Recovery

2023-12-15 · Jiajun Zhang, Yuxiang Zhang, Hongwen Zhang, Xiao Zhou 외

Accurately modeling detailed interactions between human/hand and object is an appealing yet challenging task. Current multi-view capture systems are only capable of reconstructing multiple subjects into a single, unified…

DescriptiveDisentanglementHuman-Object Interaction DetectionObject+1

Why Report Failed Interactions With Robots?! Towards Vignette-based Interaction Quality

2025-08-14 · Agnes Axelsson, Merle Reimann, Ronald Cumbal, Hannah Pelikan 외 arxiv

Although the quality of human-robot interactions has improved with the advent of LLMs, there are still various factors that cause systems to be sub-optimal when compared to human-human interactions. The nature and critic…

RFM-HRI : A Multimodal Dataset of Medical Robot Failure, User Reaction and Recovery Preferences for Item Retrieval Tasks

2026-03-05 · Yashika Batra, Giuliano Pioldi, Promise Ekpo, Arman Sayatqyzy 외 arxiv

While robots deployed in real-world environments inevitably experience interaction failures, understanding how users respond through verbal and non-verbal behaviors remains under-explored in human-robot interaction (HRI)…

Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering

2026-09-15 · Kevin Mo, Nathan Mo, Richard Zhu arxiv

Multi-hop question answering requires combining information from multiple documents to answer complex questions. These systems have grown increasingly capable, yet when they fail, the error is typically attributed to not…

Multi-hop Question Answering