paper-with-me

홈 › Papers

REPAIR-Bench: A Benchmark for Robot Error Perception And Interaction Recovery

2026-06-29 · Giuliano Pioldi, Yashika Batra, Arman Ibrayeva, Yuanchen Bai, Purnjay Maruur, Promise Ekpo, Angelique Taylor arxiv

Understanding how users perceive and respond to robot failures is essential for building robust and trustworthy robot systems. Prior work, however, (i) often treats failures as independent events, (ii) emphasizes binary failure detection, (iii) with rule-based recovery modeling. We present REPAIR-Bench, built on 214 interaction trials from 41 participants, the benchmark spans four induced failure types and provides synchronized facial action units, head pose, speech transcripts, and post-interaction affect and recovery reports. The benchmark spans three novel evaluation tasks that jointly capture the lifecycle of failure in human-robot interaction (HRI): (i) failure detection over inter-dependent interaction sessions, modeling longitudinal user adaptation across repeated failures; (ii) visual failure-type classification beyond binary success/failure formulations; and (iii) user-centered recovery prediction, inferring users' preferred recovery strategies from interaction context rather than relying on manually designed or rule-based strategies. In baseline experiments, hierarchical recurrent modeling improved failure detection over a single-session model (strict F1: 0.80 vs. 0.68), achieved a failure localization mean signed error of -0.51 s, median absolute error of 2.97 s and, for recovery prediction, a QLoRA-tuned Mistral-7B reached Hit@5=0.76 and F1@5=0.32. REPAIR-Bench provides both the HRI and Medical HRI communities with a standardized framework for (1) evaluating robot failures and (2) building transparent, adaptive, and trustworthy recovery systems.

📄 PDF Abstract BibTeX arXiv:2606.29937

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Creating and Repairing Robot Programs in Open-World Domains

2024-10-24 · Claire Schlesinger, Arjun Guha, Joydeep Biswas

Using Large Language Models (LLMs) to produce robot programs from natural language has allowed for robot systems that can complete a higher diversity of tasks. However, LLM-generated programs may be faulty, either due to…

Diversity

A Study of In-Context-Learning-Based Text-to-SQL Errors

2025-01-16 · Jiawei Shen, Chengcheng Wan, Ruoyi Qiao, Jiazhen Zou 외

Large language models (LLMs) have been adopted to perform text-to-SQL tasks, utilizing their in-context learning (ICL) capability to translate natural language questions into structured query language (SQL). However, suc…

In-Context LearningText to SQLText-To-SQL

How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks

2026-04-12 · Johin Johny Arimbur arxiv

Large language models frequently fail to produce correct code on their first attempt, yet most benchmarks evaluate them in a single-shot setting. We investigate iterative self-repair (feeding execution errors back to the…

Code Generation

Vision-Based Adaptive Robotics for Autonomous Surface Crack Repair

2024-07-23 · Joshua Genova, Eric Cabrera, Vedhus Hoskere

Surface cracks in infrastructure can lead to significant deterioration and costly maintenance if not efficiently repaired. Manual repair methods are labor-intensive, time-consuming, and imprecise and thus difficult to sc…

CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding

2026-08-01 · Aofan Liu, Jingxiang Meng, Fangxin Liu, Yongbiao Chen arxiv

Speculative decoding mitigates the latency of sequential generation in autoregressive Large Language Models (LLMs) by interleaving draft generation with target verification. However, existing parallel drafting backends o…

Mathematical Reasoning