paper-with-me

Papers

Towards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection

2026-01-30 · Xiaoxuan Guo, Yuankun Xie, Haonan Cheng, Jiayi Zhou, Jian Liu, Hengyan Huang, Long Ye, Qin Zhang arxiv

Speech deepfake detection (SDD) focuses on identifying whether a given speech signal is genuine or has been synthetically generated. Existing audio large language model (LLM)-based methods excel in content understanding; however, their predictions are often biased toward semantically correlated cues, which results in fine-grained acoustic artifacts being overlooked during the decisionmaking process. Consequently, fake speech with natural semantics can bypass detectors despite harboring subtle acoustic anomalies; this suggests that the challenge stems not from the absence of acoustic data, but from its inadequate accessibility when semantic-dominant reasoning prevails. To address this issue, we investigate SDD within the audio LLM paradigm and introduce SDD with Auditory Perception-enhanced Audio Large Language Model (SDD-APALLM), an acoustically enhanced framework designed to explicitly expose fine-grained time-frequency evidence as accessible acoustic cues. By combining raw audio with structured spectrograms, the proposed framework empowers audio LLMs to more effectively capture subtle acoustic inconsistencies without compromising their semantic understanding. Experimental results indicate consistent gains in detection accuracy and robustness, especially in cases where semantic cues are misleading. Further analysis reveals that these improvements stem from a coordinated utilization of semantic and acoustic information, as opposed to simple modality aggregation.

📄 PDF Abstract BibTeX arXiv:2601.23066

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake Detection

Similar Papers 제목 키워드 기반

EvA: An Evidence-First Audio Understanding Paradigm for LALMs

2026-03-29 · Xinyuan Xie, Shunian Chen, Zhiheng Liu, Yuhao Zhang 외 arxiv

Large Audio Language Models (LALMs) still struggle in complex acoustic scenes because they often fail to preserve task-relevant acoustic evidence before reasoning begins. We identify this error pattern as the evidence bo…

Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs

2026-04-14 · Linhao Zhang, Yuhan Song, Aiwei Liu, Chuhan Wu 외 arxiv

Recent Audio Large Language Models (AudioLLMs) exhibit a striking performance inversion: while excelling at complex reasoning tasks, they consistently underperform on fine-grained acoustic perception. We attribute this g…

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

2026-05-29 · Yuhan Song, Linhao Zhang, Aiwei Liu, Chuhan Wu 외 arxiv

Semantic speech tokenizers have become a widely used interface for Audio-LLMs, owing to their compact single-codebook design and strong linguistic alignment. However, their focus on linguistic abstraction induces acousti…

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

2026-08-27 · Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He 외 arxiv

Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on tran…

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

2026-08-03 · Fangxu Yu, Tao Feng, Dehai Min, Zinan Lin 외 hf

Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: …

Reinforcement Learning