paper-with-me

홈 › Papers

A General Framework for Learning Procedural Audio Models of Environmental Sounds

2023-03-04 · Danzel Serrano, Mark Cartwright

This paper introduces the Procedural (audio) Variational autoEncoder (ProVE) framework as a general approach to learning Procedural Audio PA models of environmental sounds with an improvement to the realism of the synthesis while maintaining provision of control over the generated sound through adjustable parameters. The framework comprises two stages: (i) Audio Class Representation, in which a latent representation space is defined by training an audio autoencoder, and (ii) Control Mapping, in which a joint function of static/temporal control variables derived from the audio and a random sample of uniform noise is learned to replace the audio encoder. We demonstrate the use of ProVE through the example of footstep sound effects on various surfaces. Our results show that ProVE models outperform both classical PA models and an adversarial-based approach in terms of sound fidelity, as measured by Fr\'echet Audio Distance (FAD), Maximum Mean Discrepancy (MMD), and subjective evaluations, making them feasible tools for sound design workflows.

📄 PDF Abstract BibTeX arXiv:2303.02396

Code (0)

등록된 구현이 없습니다.

Tasks

FAD

Similar Papers 제목 키워드 기반

EnvSDD: Benchmarking Environmental Sound Deepfake Detection

2025-05-25 · Han Yin, Yang Xiao, Rohan Kumar Das, Jisheng Bai 외

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental …

Audio Deepfake DetectionAudio GenerationBenchmarkingDeepFake Detection+1

EnvGAN: Adversarial Synthesis of Environmental Sounds for Data Augmentation

2021-04-15 · Aswathy Madhu, Suresh K

The research in Environmental Sound Classification (ESC) has been progressively growing with the emergence of deep learning algorithms. However, data scarcity poses a major hurdle for any huge advance in this domain. Dat…

Data AugmentationEnvironmental Sound ClassificationSound Classification

SeeingSounds: Learning Audio-to-Visual Alignment via Text

2025-10-10 · Simone Carnemolla, Matteo Pennisi, Chiara Russo, Simone Palazzo 외 arxiv

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on …

Image Generation

Expressive Range Characterization of Open Text-to-Audio Models

2025-10-31 · Jonathan Morse, Azadeh Naderi, Swen Gaudl, Mark Cartwright 외 arxiv

Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional content that they create (e.g., playabi…

Environmental Sound Classification

SounDiT: Geo-Contextual Soundscape-to-Landscape Generation

2025-05-19 · JunBo Wang, Haofeng Tan, Bowen Liao, Albert Jiang 외

We present a novel and practically significant problem-Geo-Contextual Soundscape-to-Landscape (GeoS2L) generation-which aims to synthesize geographically realistic landscape images from environmental soundscapes. Prior a…

Image Generation