paper-with-me

홈 › Papers

Hidden Echoes Survive Training in Audio To Audio Generative Instrument Models

2024-12-14 · Christopher J. Tralie, Matt Amery, Benjamin Douglas, Ian Utz

As generative techniques pervade the audio domain, there has been increasing interest in tracing back through these complicated models to understand how they draw on their training data to synthesize new examples, both to ensure that they use properly licensed data and also to elucidate their black box behavior. In this paper, we show that if imperceptible echoes are hidden in the training data, a wide variety of audio to audio architectures (differentiable digital signal processing (DDSP), Realtime Audio Variational autoEncoder (RAVE), and ``Dance Diffusion'') will reproduce these echoes in their outputs. Hiding a single echo is particularly robust across all architectures, but we also show promising results hiding longer time spread echo patterns for an increased information capacity. We conclude by showing that echoes make their way into fine tuned models, that they survive mixing/demixing, and that they survive pitch shift augmentation during training. Hence, this simple, classical idea in watermarking shows significant promise for tagging generative audio models.

📄 PDF Abstract BibTeX arXiv:2412.10649

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Echoes: A semantically-aligned music deepfake detection dataset

2026-03-24 · Octavian Pascu, Dan Oneata, Horia Cucu, Nicolas M. Muller arxiv

We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions. Echoes comprises 4,468 tracks (131 hours of audio) spannin…

DeepFake DetectionMusic Generation

Beyond Image to Depth: Improving Depth Prediction using Echoes

2021-03-15 · CVPR 2021 1 · Kranti Kumar Parida, Siddharth Srivastava, Gaurav Sharma

We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utili…

Depth EstimationDepth PredictionPrediction

SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset

2024-05-12 · Sushant Gautam, Mehdi Houshmand Sarkhoosh, Jan Held, Cise Midoglu 외

The application of Automatic Speech Recognition (ASR) technology in soccer offers numerous opportunities for sports analytics. Specifically, extracting audio commentaries with ASR provides valuable insights into the even…

Action SpottingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Caption Generation+3

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

2026-02-24 · Christian Simon, Masato Ishii, Wei-Yao Wang, Koichi Saito 외 arxiv

Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information. In this work, we tackle the scaling chal…

Audio Generation

Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation

2024-05-23 · Shiqi Yang, Zhi Zhong, Mengjie Zhao, Shusuke Takahashi 외

In recent years, with the realistic generation results and a wide range of personalized applications, diffusion-based generative models gain huge attention in both visual and audio generation areas. Compared to the consi…

Audio GenerationDenoisingLanguage ModelingLanguage Modelling+1