paper-with-me

홈 › Papers

SARA: Stress Test Reasoning in Audio Deepfake Detection

2026-01-07 · Binh Nguyen, Charles Fleming, Thai Le arxiv

Audio Language Models (ALMs) offer a promising shift towards explainable audio deepfake detections (ADD), moving beyond \textit{black-box} classifiers by providing transparency to their predictions via reasoning traces. However, such reasoning may not support the model predictions, reflecting poor coherence, or, worse, may rationalize incorrect predictions with plausible but misleading explanation. Moreover, the behavior of ALM reasoning under adversarial attacks remains under-explored, raising questions about the practical reliability of such explanation capabilities. To address this gap, this study introduces \textbf{SARA} (\textbf{S}hift \textbf{A}nalysis of \textbf{R}easoning in \textbf{A}udio), a diagnostic framework that evaluates ALM reasoning across three dimensions: acoustic perception, reasoning-verdict coherence and dissonance. We test five open-source ALMs against both acoustic and linguistic adversarial attacks. We show that acoustic attacks significantly degrade reasoning-verdict coherence (average decrease of 14.20\%), frequently inducing internal logical conflicts. Conversely, linguistic attacks achieve higher attack success rates while maintaining reasoning coherence. We further demonstrate that the textual coherence of generated reasoning traces also serves as a latent indicator of adversarial inputs, enabling effective detection of perturbed audio (0.78 in F1) \textit{without accessing the raw acoustic signal}. These findings suggest that reasoning traces provide diagnostic utility that persists even when final classification outputs are compromised.

📄 PDF Abstract BibTeX arXiv:2601.03615

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Deepfake Detection

Similar Papers 제목 키워드 기반

Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?

2024-08-20 · Yuankun Xie, Chenxu Xiong, Xiaopeng Wang, Zhiyong Wang 외

Currently, Audio Language Models (ALMs) are rapidly advancing due to the developments in large language models and audio neural codecs. These ALMs have significantly lowered the barrier to creating deepfake audio, genera…

Audio Deepfake DetectionDeepFake DetectionFace Swapping

Detecting music deepfakes is easy but actually hard

2024-05-07 · Darius Afchar, Gabriel Meseguer-Brocal, Romain Hennequin

In the face of a new era of generative models, the detection of artificially generated content has become a matter of utmost importance. The ability to create credible minute-long music deepfakes in a few seconds on user…

DeepFake DetectionFace Swapping

XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark

2025-05-31 · Ioan-Paul Ciobanu, Andrei-Iulian Hiji, Nicolae-Catalin Ristea, Paul Irofti 외

Recent advances in audio generation led to an increasing number of deepfakes, making the general public more vulnerable to financial scams, identity theft, and misinformation. Audio deepfake detectors promise to alleviat…

Audio GenerationFace SwappingMisinformation

StressTest: Can YOUR Speech LM Handle the Stress?

2025-05-28 · Iddo Yosha, Gallil Maimon, Yossi Adi

Sentence stress refers to emphasis, placed on specific words within a spoken utterance to highlight or contrast an idea, or to introduce new information. It is often used to imply an underlying intention that is not expl…

Question AnsweringSentenceSynthetic Data Generation

Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection

2026-01-02 · Akanksha Chuchra, Shukesh Reddy, Sudeepta Mishra, Abhijit Das 외 arxiv

While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored.…

Audio Deepfake Detection