paper-with-me

홈 › Papers

Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions

2026-04-19 · Dongwook Lee, Eunwoo Song, Che Hyun Lee, Heeseung Kim, Sungroh Yoon arxiv

While recent Spoken Language Models (SLMs) have been actively deployed in real-world scenarios, they lack the capability to discern Third-Party Interruptions (TPI) from the primary user's ongoing flow, leaving them vulnerable to contextual failures. To bridge this gap, we introduce TPI-Train, a dataset of 88K instances designed with speaker-aware hard negatives to enforce acoustic cue prioritization for interruption handling, and TPI-Bench, a comprehensive evaluation framework designed to rigorously measure the interruption-handling strategy and precise speaker discrimination in deceptive contexts. Experiments demonstrate that our dataset design mitigates semantic shortcut learning-a critical pitfall where models exploit semantic context while neglecting acoustic signals essential for discerning speaker changes. We believe our work establishes a foundational resource for overcoming text-dominated unimodal reliance in SLMs, paving the way for more robust multi-party spoken interaction. The code for the framework is publicly available at https://tpi-va.github.io

📄 PDF Abstract BibTeX arXiv:2604.17358

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing

2025-09-26 · Ke Wang, Houxing Ren, Zimu Lu, Mingjie Zhan 외 arxiv

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabili…

Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants

2024-05-14 · Chloé Sekkat, Fanny Leroy, Salima Mdhaffar, Blake Perry Smith 외

Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets w…

Automatic Speech RecognitionDiversityspeech-recognitionSpeech Recognition+1

VoiceAgentBench: Are Voice Assistants ready for agentic tasks?

2025-10-09 · Dhruv Jain, Harshit Shukla, Gautam Rajeev, Ashish Kulkarni 외 arxiv

Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks largely focus on isolated capabilities su…

Adversarial RobustnessQuestion AnsweringVoice Conversion

Evaluating Personal Assistants on Mobile devices

2017-06-14 · Kiseleva Julia, de Rijke Maarten

The iPhone was introduced only a decade ago in 2007 but has fundamentally changed the way we interact with online information. Mobile devices differ radically from classic command-based and point-and-click user interface…

Learning to Rank Intents in Voice Assistants

2020-04-30 · Raviteja Anantha, Srinivas Chappidi, William Dawoodi

Voice Assistants aim to fulfill user requests by choosing the best intent from multiple options generated by its Automated Speech Recognition and Natural Language Understanding sub-systems. However, voice assistants do n…

DenoisingLearning-To-RankNatural Language Understandingspeech-recognition+1