paper-with-me

홈 › Papers

HearSay Benchmark: Do Audio LLMs Leak What They Hear?

2026-01-07 · Jin Wang, Liang Lin, Kaiwen Luo, Weiliu Wang, Yitian Chen, Moayad Aloqaily, Xuehai Tang, Zhenhong Zhou, Kun Wang, Li Sun, Qingsong Wen arxiv

While Audio Large Language Models (ALLMs) have achieved remarkable progress in understanding and generation, their potential privacy implications remain largely unexplored. This paper takes the first step to investigate whether ALLMs inadvertently leak user privacy solely through acoustic voiceprints and introduces $\textit{HearSay}$, a comprehensive benchmark constructed from over 22,000 real-world audio clips. To ensure data quality, the benchmark is meticulously curated through a rigorous pipeline involving automated profiling and human verification, guaranteeing that all privacy labels are grounded in factual records. Extensive experiments on $\textit{HearSay}$ yield three critical findings: $\textbf{Significant Privacy Leakage}$: ALLMs inherently extract private attributes from voiceprints, reaching 92.89% accuracy on gender and effectively profiling social attributes. $\textbf{Insufficient Safety Mechanisms}$: Alarmingly, existing safeguards are severely inadequate; most models fail to refuse privacy-intruding requests, exhibiting near-zero refusal rates for physiological traits. $\textbf{Reasoning Amplifies Risk}$: Chain-of-Thought (CoT) reasoning exacerbates privacy risks in capable models by uncovering deeper acoustic correlations. These findings expose critical vulnerabilities in ALLMs, underscoring the urgent need for targeted privacy alignment. The codes and dataset are available at https://github.com/JinWang79/HearSay_Benchmark

📄 PDF Abstract BibTeX arXiv:2601.03783

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact

2026-08-04 · Alex Kwon arxiv

AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, who said it, how sure they were, when it …

What Are They Doing? Joint Audio-Speech Co-Reasoning

2024-09-22 · Yingzhi Wang, Pooneh Mousavi, Artem Ploujnikov, Mirco Ravanelli

In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Auditory Large Language Models (ALLMs) have ma…

Protecting Bystander Privacy via Selective Hearing in Audio LLMs

2025-12-06 · Xiao Zhan, Guangzhi Sun, Jose Such, Phil Woodland arxiv

Audio Large language models (LLMs) are increasingly deployed in the real world, where they inevitably capture speech from unintended nearby bystanders, raising privacy risks that existing benchmarks and defences did not …

Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery

2026-06-03 · Arquimedes Canedo, Grama Chethan arxiv

When an AI agent calls an API and hits a validation error, it needs more than what went wrong -- it needs what to do next. A self-reflective API returns, on validation failure, a machine-readable recovery\_feedback.sugge…

LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks

2025-02-10 · Xin Zhou, Martin Weyssow, Ratnadira Widyasari, Ting Zhang 외

Large Language Models (LLMs) are widely utilized in software engineering (SE) tasks, such as code generation and automated program repair. However, their reliance on extensive and often undisclosed pre-training datasets …

Code GenerationProgram Repair