paper-with-me

홈 › Papers

Generation or Replication: Auscultating Audio Latent Diffusion Models

2023-10-16 · Dimitrios Bralios, Gordon Wichern, François G. Germain, Zexu Pan, Sameer Khurana, Chiori Hori, Jonathan Le Roux

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt at understanding the inner workings of audio latent diffusion models by investigating how their audio outputs compare with the training data, similar to how a doctor auscultates a patient by listening to the sounds of their organs. Using text-to-audio latent diffusion models trained on the AudioCaps dataset, we systematically analyze memorization behavior as a function of training set size. We also evaluate different retrieval metrics for evidence of training data memorization, finding the similarity between mel spectrograms to be more robust in detecting matches than learned embedding vectors. In the process of analyzing memorization in audio latent diffusion models, we also discover a large amount of duplicated audio clips within the AudioCaps database.

📄 PDF Abstract BibTeX arXiv:2310.10604

Code (0)

등록된 구현이 없습니다.

Tasks

AudioCapsMemorizationRetrieval

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance

2025-09-18 · Francisco Messina, Francesca Ronchini, Luca Comanducci, Paolo Bestagini 외 arxiv

A persistent challenge in generative audio models is data replication, where the model unintentionally generates parts of its training data during inference. In this work, we address this issue in text-to-audio diffusion…

Audio Generation

Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

2024-02-27 · CVPR 2024 1 · Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang 외

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the tech…

Audio GenerationDenoising

Stable Audio 3

2026-05-18 · Zach Evans, Julian D. Parker, Matthew Rice, CJ Carr 외 arxiv

Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of audio, variable-length generations are …

Audio Generation

LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation

2024-06-12 · Wenhao Guan, Kaidi Wang, Wangjin Zhou, Yang Wang 외

Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusion models still needs improvement. And th…

Audio Generation

Latent CLAP Loss for Better Foley Sound Synthesis

2024-03-18 · Tornike Karchkhadze, Hassan Salami Kavaki, Mohammad Rasool Izadi, Bryce Irvin 외

Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimodal text-audio representation models, su…

FAD