paper-with-me

홈 › Papers

IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

2026-06-17 · Ahmad Salimi, Wentao Ma, Yuzhi Tang, Dongming Shen, Mu Li, Alex Smola arxiv

Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user interruptions while maintaining progress through multi-step procedures. Existing benchmarks for speech-capable models focus on the timing of interruptions: barge-in detection, endpointing, and turn-taking dynamics. They leave unmeasured what happens after the interruption: does the agent resume the workflow at the correct step? Does it address the user's interjection? Does it avoid re-delivering content the user already heard? We introduce IHBench (Interruption Handling Benchmark), a benchmark that evaluates post-interruption recovery in voice agents executing state-machine-driven workflows across 10 enterprise domains. Six interruption types are injected at controlled points mid-utterance, with per-interruption evaluation rubrics generated alongside the data. Each interruption is scored on two axes: task fulfillment and recovery quality. We evaluate 27 audio-language model configurations from OpenAI, Google, and the open-weight community. Models vary widely, and recovery quality depends strongly on the interruption type. Across our experiments, closed-weight models are consistently more robust to interruptions than open-weight ones: they win far more often on task fulfillment, degrade roughly 3.3x more slowly as conversations grow longer, and show no audio-versus-text modality gap, whereas the open-weight models lose ground on all three. A human study validates the LLM judge against human annotators, and a cross-benchmark analysis against AudioMultiChallenge indicates that recovery quality is a largely distinct capability axis.

📄 PDF Abstract BibTeX arXiv:2606.19595

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions

2026-04-08 · Smit Nautambhai Modi, Gandharv Mahajan, Marc Wetter, Randall Welles arxiv

Real-time voice assistants must revise task state when users interrupt mid-response, but existing spoken-dialog benchmarks largely evaluate turn-based interaction and miss this failure mode. We introduce EchoChain, a con…

Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions

2026-04-19 · Dongwook Lee, Eunwoo Song, Che Hyun Lee, Heeseung Kim 외 arxiv

While recent Spoken Language Models (SLMs) have been actively deployed in real-world scenarios, they lack the capability to discern Third-Party Interruptions (TPI) from the primary user's ongoing flow, leaving them vulne…

MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models

2025-08-01 · Jiale Li, Mingrui Wu, Zixiang Jin, Hao Chen 외 arxiv

Despite growing interest in hallucination in Multimodal Large Language Models, existing studies primarily focus on single-image settings, leaving hallucination in multi-image scenarios largely unexplored. To address this…

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

2026-08-13 · Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor, Shehzeen Hussain 외 arxiv

Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-spe…

Speech Synthesis

Interruption-Aware Cooperative Perception for V2X Communication-Aided Autonomous Driving

2023-04-24 · Shunli Ren, Zixing Lei, Zi Wang, Mehrdad Dianati 외

Cooperative perception can significantly improve the perception performance of autonomous vehicles beyond the limited perception ability of individual vehicles by exchanging information with neighbor agents through V2X c…

Autonomous DrivingAutonomous VehiclesKnowledge Distillation