paper-with-me

홈 › Papers

Pardon? Evaluating Conversational Repair in Large Audio-Language Models

2026-01-19 · Shuanghong Huang, Jinlei Xu, Youchao Zhou, Yanghao Zhou, Xuan Zhao, Chong Feng, Wenxuan Zhang arxiv

Large Audio-Language Models (LALMs) have demonstrated strong performance in spoken question answering (QA), with existing evaluations primarily focusing on answer accuracy and robustness to acoustic perturbations. However, such evaluations implicitly assume that spoken inputs remain semantically answerable, an assumption that often fails in real-world interaction when essential information is missing. In this work, we introduce a repair-aware evaluation setting that explicitly distinguishes between answerable and unanswerable audio inputs. We define answerability as a property of the input itself and construct paired evaluation conditions using a semantic-acoustic masking protocol. Based on this setting, we propose the Evaluability Awareness and Repair (EAR) score, a non-compensatory metric that jointly evaluates task competence under answerable conditions and repair behavior under unanswerable conditions. Experiments on two spoken QA benchmarks across diverse LALMs reveal a consistent gap between answer accuracy and conversational reliability: while many models perform well when inputs are answerable, most fail to recognize semantic unanswerability and initiate appropriate conversational repair. These findings expose a limitation of prevailing accuracy-centric evaluation practices and motivate reliability assessments that treat unanswerable inputs as cues for repair and continued interaction.

📄 PDF Abstract BibTeX arXiv:2601.12973

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

2024-10-06 · Anton Cheshkov, Pavel Zadorozhny, Rodion Levichev, Evgeny Maslov 외

Automatic program repair at project level may open yet to be seen opportunities in various fields of human activity. Since the SWE-Bench challenge was presented, we have seen numerous of solutions. Patch generation is a …

Program Repairvalid

No that's not what I meant: Handling Third Position Repair in Conversational Question Answering

2023-07-31 · Vevake Balaraman, Arash Eshghi, Ioannis Konstas, Ioannis Papaioannou

The ability to handle miscommunication is crucial to robust and faithful conversational AI. People usually deal with miscommunication immediately as they detect it, using highly systematic interactional mechanisms called…

Conversational Question AnsweringPositionQuestion Answering

NC-Bench: An LLM Benchmark for Evaluating Conversational Competence

2026-01-10 · Robert J. Moore, Sungeun An, Farhan Ahmed, Jay Pankaj Gala arxiv

The Natural Conversation Benchmark (NC-Bench) introduces a new approach to evaluating the general conversational competence of large language models (LLMs). Unlike prior benchmarks that focus on the content of model beha…

"Mm, Wat?" Detecting Other-initiated Repair Requests in Dialogue

2025-10-28 · Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chloe Clavel arxiv

Maintaining mutual understanding is a key component in human-human conversation to avoid conversation breakdowns, in which repair, particularly Other-Initiated Repair (OIR, when one speaker signals trouble and prompts th…

Pardon the Interruption: Automatic Analysis of Gender and Competitive Turn-Taking in United States Supreme Court Hearings

2019-08-01 · WS 2019 8 · Haley Lepp

The United States Supreme Court plays a key role in defining the legal basis for gender discrimination throughout the country, yet there are few checks on gender bias within the court itself. In conversational turn-takin…