paper-with-me

홈 › Papers

WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection

2025-10-06 · Xi Xuan, Xuechen Liu, Wenxin Zhang, Yi-Cheng Lin, Xiaojian Lin, Tomi Kinnunen arxiv

Modern front-end design for speech deepfake detection relies on full fine-tuning of large pre-trained models like XLSR. However, this approach is not parameter-efficient and may lead to suboptimal generalization to realistic, in-the-wild data types. To address these limitations, we introduce a new family of parameter-efficient front-ends that fuse prompt-tuning with classical signal processing transforms. These include FourierPT-XLSR, which uses the Fourier Transform, and two variants based on the Wavelet Transform: WSPT-XLSR and Partial-WSPT-XLSR. We further propose WaveSP-Net, a novel architecture combining a Partial-WSPT-XLSR front-end and a bidirectional Mamba-based back-end. This design injects multi-resolution features into the prompt embeddings, which enhances the localization of subtle synthetic artifacts without altering the frozen XLSR parameters. Experimental results demonstrate that WaveSP-Net outperforms several state-of-the-art models on two new and challenging benchmarks, Deepfake-Eval-2024 and SpoofCeleb, with low trainable parameters and notable performance gains. The code and models are available at https://github.com/xxuan-acoustics/WaveSP-Net.

📄 PDF Abstract BibTeX arXiv:2510.05305

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake Detection

Similar Papers 제목 키워드 기반

Wavesplit: End-to-End Speech Separation by Speaker Clustering

2020-02-20 · Neil Zeghidour, David Grangier

We introduce Wavesplit, an end-to-end source separation system. From a single mixture, the model infers a representation for each source and then estimates each source signal given the inferred representations. The model…

ClusteringData AugmentationSpeech Separation

Fully Learnable Deep Wavelet Transform for Unsupervised Monitoring of High-Frequency Time Series

2021-05-03 · Gabriel Michau, Gaetan Frusque, Olga Fink

High-Frequency (HF) signals are ubiquitous in the industrial world and are of great use for monitoring of industrial assets. Most deep learning tools are designed for inputs of fixed and/or very limited size and many suc…

Deep LearningDenoisingTime SeriesTime Series Analysis

PointWavelet: Learning in Spectral Domain for 3D Point Cloud Analysis

2023-02-10 · Cheng Wen, Jianzhi Long, Baosheng Yu, DaCheng Tao

With recent success of deep learning in 2D visual recognition, deep learning-based 3D point cloud analysis has received increasing attention from the community, especially due to the rapid development of autonomous drivi…

Autonomous DrivingDeep LearningPoint Cloud Classification

WaveRNet: Wavelet-Guided Frequency Learning for Multi-Source Domain-Generalized Retinal Vessel Segmentation

2026-01-09 · Chanchan Wang, Yuanfang Wang, Qing Xu, Guanxin Chen arxiv

Domain-generalized retinal vessel segmentation is critical for automated ophthalmic diagnosis, yet faces significant challenges from domain shift induced by non-uniform illumination and varying contrast, compounded by th…

Retinal Vessel Segmentation

Sparse wavelet-based solutions for the M/EEG inverse problem

2023-06-27 · Samy Mokhtari, Jean-Michel Badier, Christian G. Bénar, Bruno Torrésani

This paper is concerned with variational and Bayesian approaches to neuro-electromagnetic inverse problems (EEG and MEG). The strong indeterminacy of these problems is tackled by introducing sparsity inducing regularizat…

Data CompressionEEGregression