paper-with-me

홈 › Papers

$\mathcal{A}LLM4ADD$: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection

2025-05-16 · Hao Gu, Jiangyan Yi, Chenglong Wang, JianHua Tao, Zheng Lian, Jiayi He, Yong Ren, Yujie Chen, Zhengqi Wen

Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large language models (ALLMs) have made significant progress in various audio processing tasks, a heuristic question arises: Can ALLMs be leveraged to solve ADD?. In this paper, we first conduct a comprehensive zero-shot evaluation of ALLMs on ADD, revealing their ineffectiveness in detecting fake audio. To enhance their performance, we propose $\mathcal{A}LLM4ADD$, an ALLM-driven framework for ADD. Specifically, we reformulate ADD task as an audio question answering problem, prompting the model with the question: "Is this audio fake or real?". We then perform supervised fine-tuning to enable the ALLM to assess the authenticity of query audio. Extensive experiments are conducted to demonstrate that our ALLM-based method can achieve superior performance in fake audio detection, particularly in data-scarce scenarios. As a pioneering study, we anticipate that this work will inspire the research community to leverage ALLMs to develop more effective ADD systems.

📄 PDF Abstract BibTeX arXiv:2505.11079

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Deepfake DetectionAudio Question AnsweringDeepFake DetectionFace SwappingQuestion Answering

Similar Papers 제목 키워드 기반

Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation

2026-01-20 · Arjun Chandra, Kevin Miller, Venkatesh Ravichandran, Constantinos Papayiannis 외 arxiv

Large Language Model (LLM) judges exhibit strong reasoning capabilities but are limited to textual content. This leaves current automatic Speech-to-Speech (S2S) evaluation methods reliant on opaque and expensive Audio La…

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio

2026-06-22 · Ran Piao, Tsai-Ning Wang, Martijn den Dekker, Linda Moonen 외 arxiv

Clinical audio diagnosis in low-resource settings requires models that identify conditions from minimal examples without large annotated corpora. We propose Federated Self-Contextualization (FSC), a multimodal language m…

Multimodal Reasoning

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

2023-01-30 · Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 외

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high…

Audio GenerationText-to-Video GenerationVideo Generation

Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off

2026-02-27 · Longyin Zhang, Shuo Sun, Yingxu He, Won Cheng Yi Lewis 외 arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) pursue omni-perception capabilities, yet integrating robust sensory grounding with complex reasoning remains a challenge, particularly for underrepresented …

Defending against Adversarial Audio via Diffusion Model

2023-03-02 · Shutong Wu, Jiongxiao Wang, Wei Ping, Weili Nie 외

Deep learning models have been widely used in commercial acoustic systems in recent years. However, adversarial audio examples can cause abnormal behaviors for those acoustic systems, while being hard for humans to perce…

Adversarial Purificationmodel