paper-with-me

홈 › Papers

SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection

2026-03-21 · Kyudan Jung, Jihwan Kim, Minwoo Lee, Soyoon Kim, Jeonghoon Kim, Jaegul Choo, Cheonbok Park arxiv

Recent advancements in text-to-speech technologies enable generating high-fidelity synthetic speech nearly indistinguishable from real human voices. While recent studies show the efficacy of self-supervised learning-based speech encoders for deepfake detection, these models struggle to generalize across unseen speakers. Our quantitative analysis suggests these encoder representations are substantially influenced by speaker information, causing detectors to exploit speaker-specific correlations rather than artifact-related cues. We call this phenomenon speaker entanglement. To mitigate this reliance, we introduce SNAP, a speaker-nulling framework. We estimate a speaker subspace and apply orthogonal projection to suppress speaker-dependent components, isolating synthesis artifacts within the residual features. By reducing speaker entanglement, SNAP encourages detectors to focus on artifact-related patterns, leading to state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2603.20686

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningDeepFake Detection

Similar Papers 제목 키워드 기반

SSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling

2026-02-01 · Yochai Yemini, Yoav Ellinson, Rami Ben-Ari, Sharon Gannot 외 arxiv

This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on generative inverse sampling, where we mod…

Speech Separation

A Spoofing Benchmark for the 2018 Voice Conversion Challenge: Leveraging from Spoofing Countermeasures for Speech Artifact Assessment

2018-04-23 · Tomi Kinnunen, Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda 외

Voice conversion (VC) aims at conversion of speaker characteristic without altering content. Due to training data limitations and modeling imperfections, it is difficult to achieve believable speaker mimicry without intr…

BenchmarkingSpeaker VerificationVoice Conversion

Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech Recognition

2022-05-09 · Catalin Zorila, Rama Doddipatla

Improving the accuracy of single-channel automatic speech recognition (ASR) in noisy conditions is challenging. Strong speech enhancement front-ends are available, however, they typically require that the ASR model is re…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

Single-Channel Speech Separation with Auxiliary Speaker Embeddings

2019-06-24 · Shuo Liu, Gil Keren, Björn Schuller

We present a novel source separation model to decompose asingle-channel speech signal into two speech segments belonging to two different speakers. The proposed model is a neural network based on residual blocks, and use…

Speech Separation

Translatotron 2: High-quality direct speech-to-speech translation with voice preservation

2021-07-19 · Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, Roi Pomerantz

We present Translatotron 2, a neural direct speech-to-speech translation model that can be trained end-to-end. Translatotron 2 consists of a speech encoder, a linguistic decoder, an acoustic synthesizer, and a single att…

Data AugmentationDecoderSpeech-to-Speech TranslationTranslation+1