paper-with-me

Papers

WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables

2025-12-25 · Zhaojiang Lin, Yong Xu, Kai Sun, Jing Zheng, Yin Huang, Surya Teja Appini, Krish Narang, Renjie Tao, Ishan Kapil Jain, Siddhant Arora, Ruizhi Li, Yiteng Huang, Kaushik Patnaik, Wenfang Xu, Suwon Shon, Yue Liu, Ahmed A Aly, Anuj Kumar, Florian Metze, Xin Luna Dong arxiv

Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research.

📄 PDF Abstract BibTeX arXiv:2601.02391

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

2025-08-06 · Liang Xu, Chengqun Yang, Zili Lin, Fei Xu 외 arxiv

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist intera…

EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

2026-08-26 · Yu-Chien Tang, Yu-Hsiang Liu, An-Zi Yen arxiv

VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation…

VoiceBench: Benchmarking LLM-Based Voice Assistants

2024-10-22 · Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao 외

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingGeneral Knowledge+2

Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges

2025-11-17 · Junlong Li, Huaiyuan Xu, Sijie Cheng, Kejun Wu 외 arxiv

Driven by recent advances in vision-language models (VLMs) and egocentric perception research, the emerging topic of an egocentric procedural AI assistant (EgoProceAssist) is introduced to step-by-step support daily proc…

Question Answering

VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing

2025-09-26 · Ke Wang, Houxing Ren, Zimu Lu, Mingjie Zhan 외 arxiv

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabili…