paper-with-me

홈 › Papers

VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models

2026-01-27 · Yuxiang Wang, Hongyu Liu, Dekun Chen, Xueyao Zhang, Zhizheng Wu arxiv

As Speech Language Models (SLMs) transition from personal devices to shared, multi-user environments such as smart homes, a new challenge emerges: the model is expected to distinguish between users to manage information flow appropriately. Without this capability, an SLM could reveal one user's confidential schedule to another, a privacy failure we term interactional privacy. Thus, the ability to generate speaker-aware responses becomes essential for SLM safe deployment. Current SLM benchmarks test dialogue ability but overlook speaker identity. Multi-speaker benchmarks check who said what without assessing whether SLMs adapt their responses. Privacy benchmarks focus on globally sensitive data (e.g., bank passwords) while neglecting contextual privacy-sensitive information (e.g., a user's private appointment). To address this gap, we introduce VoxPrivacy, the first benchmark designed to evaluate interactional privacy in SLMs. VoxPrivacy spans three tiers of increasing difficulty, from following direct secrecy commands to proactively protecting privacy. Our evaluation of nine SLMs on a 32-hour bilingual dataset reveals a widespread vulnerability: most open-source models perform close to random chance (around 50% accuracy) on conditional privacy decisions, while even strong closed-source systems fall short on proactive privacy inference. We further validate these findings on Real-VoxPrivacy, a human-recorded subset, confirming that failures observed on synthetic data persist in real speech. Finally, we demonstrate a viable path forward: by fine-tuning on a new 4,000-hour training set, we improve privacy-preserving abilities while maintaining robustness. To support future work, we release the VoxPrivacy benchmark, the large-scale training set, and the fine-tuned model to foster the development of safer and more context-aware SLMs.

📄 PDF Abstract BibTeX arXiv:2601.19956

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

2026-07-16 · David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer 외 arxiv

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information …

Speech Recognition

TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios

2025-07-24 · Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang 외 arxiv

Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited…

From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines

2025-12-12 · Tittaya Mairittha, Tanakon Sawanglok, Panuwit Raden, Jirapast Buntub 외 arxiv

While voice-based AI systems have achieved remarkable generative capabilities, their interactions often feel conversationally broken. This paper examines the interactional friction that emerges in modular Speech-to-Speec…

Recurrent Modeling of Interaction Context for Collective Activity Recognition

2017-07-01 · CVPR 2017 7 · Minsi Wang, Bingbing Ni, Xiaokang Yang

Modeling of high order interactional context, e.g., group interaction, lies in the central of collective/group activity recognition. However, most of the previous activity recognition methods do not offer a flexible and…

Activity RecognitionDescriptiveGroup Activity Recognition

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

2026-04-20 · Suhyun Lee, Palakorn Achananuparp, Neemesh Yadav, Ee-Peng Lim 외 arxiv

Large language models (LLMs) are increasingly explored as scalable tools for mental health counseling, yet evaluating their safety remains challenging due to the interactional and context-dependent nature of clinical har…