paper-with-me

홈 › Papers

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

2026-07-29 · Haotian Mo, Jie Liu, Siqi Shen, Songzhu Mei, Xinhai Chen, Xiangyang Wang, Yigui Feng, Shuai Li, Gencheng Liu, Keqi Yang, Qinglin Wang arxiv

Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe. Reconstructions at masking ratios 0.5, 0.75, and 0.9 yield explicit multi-ratio residual maps. Because these residuals are domain sensitive, our audio-anchored detector passes the projected frozen-WavLM auditory representation into the fusion sum without gate-based attenuation and uses residuals only as a scalar-gated additive correction. The pre-specified seed-42 run obtains 6.5442% EER / 0.18456 min-DCF on ASVspoof 5 Eval and 13.8372% / 0.36921 on ITW Full; three-seed means are 6.8885 (0.3308)% and 15.3328 (2.0719)%. The latter is below a separately optimized WavLM-ResNet18 reference under both supervision settings. Auxiliary supervision raises dynamic competitive fusion from 18.4007% to 25.2968% mean ITW EER, worsening all three seeds. The results support reconstruction residuals as complementary evidence and motivate a non-competitive auditory path for ASVspoof 5-to-ITW transfer, without claiming a componentwise causal ablation of anchoring alone.

📄 PDF Abstract BibTeX arXiv:2607.26472

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Deepfake Detection

Similar Papers 제목 키워드 기반

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

2026-09-24 · Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li 외 hf

Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. Whil…

Reinforcement LearningVideo Generation

Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation

2026-01-21 · Yifei Liu, Changxing Ding, Ling Guo, Huaiguang Jiang 외 arxiv

Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabilities and flexibility. However, current motion diffusion models face two …

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

2026-05-17 · Yuheng Chen, Qingdong He, Teng Hu, Yuji Wang 외 arxiv

The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides, achieving cohesive multimodal customization for the simultaneous pr…

Video Generation

FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

2024-10-16 · Huadai Liu, Jialei Wang, Rongjie Huang, Yang Liu 외

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While…

Audio GenerationGPU

Anchored Diffusion Language Model

2025-05-24 · Litu Rout, Constantine Caramanis, Sanjay Shakkottai

Diffusion Language Models (DLMs) promise parallel generation and bidirectional context, yet they underperform autoregressive (AR) models in both likelihood modeling and generated text quality. We identify that this perfo…

Language ModelingLanguage ModellingMathmodel+1