paper-with-me

Papers

Synthetic Speech Source Tracing using Metric Learning

2025-06-03 · Dimitrios Koutsianos, Stavros Zacharopoulos, Yannis Panagakis, Themos Stafylakis

This paper addresses source tracing in synthetic speech-identifying generative systems behind manipulated audio via speaker recognition-inspired pipelines. While prior work focuses on spoofing detection, source tracing lacks robust solutions. We evaluate two approaches: classification-based and metric-learning. We tested our methods on the MLAADv5 benchmark using ResNet and self-supervised learning (SSL) backbones. The results show that ResNet achieves competitive performance with the metric learning approach, matching and even exceeding SSL-based systems. Our work demonstrates ResNet's viability for source tracing while underscoring the need to optimize SSL representations for this task. Our work bridges speaker recognition methodologies with audio forensic challenges, offering new directions for combating synthetic media manipulation.

📄 PDF Abstract BibTeX arXiv:2506.02590

Code (0)

등록된 구현이 없습니다.

Tasks

Metric LearningSelf-Supervised LearningSpeaker Recognition

Methods 이 논문이 사용한 방법론

Average Pooling 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Kaiming Initialization 설명 없음
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…

Similar Papers 제목 키워드 기반

Synthetic Wave-Geometric Impulse Responses for Improved Speech Dereverberation

2022-12-10 · Rohith Aralikatti, Zhenyu Tang, Dinesh Manocha

We present a novel approach to improve the performance of learning-based speech dereverberation using accurate synthetic datasets. Our approach is designed to recover the reverb-free signal from a reverberant speech sign…

Speech Dereverberation

Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion

2025-06-02 · Ajinkya Kulkarni, Sandipana Dowerah, Tanel Alumae, Mathew Magimai. -Doss

Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source system is equally crucial. This work prop…

Face SwappingMetric Learning

Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations

2025-06-01 · Girish, Mohd Mujtaba Akhtar, Orchid Chetia Phukan, Drishti Singh 외

In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and intonation--into their synthesized speech,…

Emotion RecognitionRhythmSpeaker RecognitionSpeech Emotion Recognition+1

Towards Generalized Source Tracing for Codec-Based Deepfake Speech

2025-06-08 · Xuanjun Chen, I-Ming Lin, Lin Zhang, Haibin Wu 외

Recent attempts at source tracing for codec-based deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tr…

Face Swapping

Open-Set Source Tracing as Compositional Factors via Structured Prototypes

2026-07-03 · Santiago Rubio, Antonio Almudévar, Antonio Miguel, Eduardo Lleida 외 arxiv

Recent research expands beyond binary anti-spoofing with the emergence of Source Tracing, the task of identifying the specific generative origins of synthetic speech. However, current research often equates a "source" wi…