paper-with-me

Papers

Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers

2026-01-06 · Mikhail Silaev, Konstantinos Drossos, Tuomas Virtanen arxiv

Generative adversarial networks (GANs) and diffusion models have recently achieved state-of-the-art performance in audio super-resolution (ADSR), producing perceptually convincing wideband audio from narrowband inputs. However, existing evaluations primarily rely on signal-level or perceptual metrics, leaving open the question of how closely the distributions of synthetic super-resolved and real wideband audio match. Here we address this problem by analyzing the separability of real and super-resolved audio in various embedding spaces. We consider both middle-band ($4\to 16$~kHz) and full-band ($16\to 48$~kHz) upsampling tasks for speech and music, training linear classifiers to distinguish real from synthetic samples based on multiple types of audio embeddings. Comparisons with objective metrics and subjective listening tests reveal that embedding-based classifiers achieve near-perfect separation, even when the generated audio attains high perceptual quality and state-of-the-art metric scores. This behavior is consistent across datasets and models, including recent diffusion-based approaches, highlighting a persistent gap between perceptual quality and true distributional fidelity in ADSR models.

📄 PDF Abstract BibTeX arXiv:2601.03443

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Super-Resolution

Similar Papers 제목 키워드 기반

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

2026-04-30 · Sharayu Nilesh Deshmukh, Kailash A. Hambarde, Joana C. Costa, Hugo Proença 외 arxiv

Current DeepFake detection scenarios are mostly binary, yet data manipulation can vary across audio, video, or both, whose variability is not captured in binary settings. Four-class audio-visual formulations address this…

DeepFake Detection

Annotation-free Automatic Music Transcription with Scalable Synthetic Data and Adversarial Domain Confusion

2023-12-16 · Gakusei Sato, Taketo Akama

Automatic Music Transcription (AMT) is a vital technology in the field of music information processing. Despite recent enhancements in performance due to machine learning techniques, current methods typically attain high…

Music Transcription

Pre-training with Synthetic Patterns for Audio

2024-10-01 · Yuchi Ishikawa, Tatsuya Komatsu, Yoshimitsu Aoki

In this paper, we propose to pre-train audio encoders using synthetic patterns instead of real audio data. Our proposed framework consists of two key elements. The first one is Masked Autoencoder (MAE), a self-supervised…

Self-Supervised Learning

Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy

2024-06-05 · Yuankun Xie, Ruibo Fu, Zhengqi Wen, Zhiyong Wang 외

With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID) categories. However, the rapid evolution…

Audio Deepfake DetectionDeepFake DetectionFace Swapping

Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics

2025-08-01 · Tom Or, Omri Azencot arxiv

Generative models achieve remarkable results in multiple data domains, including images and texts, among other examples. Unfortunately, malicious users exploit synthetic media for spreading misinformation and disseminati…