paper-with-me

Papers

VIFS: An End-to-End Variational Inference for Foley Sound Synthesis

2023-06-08 · Junhyeok Lee, Hyeonuk Nam, Yong-Hwa Park

The goal of DCASE 2023 Challenge Task 7 is to generate various sound clips for Foley sound synthesis (FSS) by "category-to-sound" approach. "Category" is expressed by a single index while corresponding "sound" covers diverse and different sound examples. To generate diverse sounds for a given category, we adopt VITS, a text-to-speech (TTS) model with variational inference. In addition, we apply various techniques from speech synthesis including PhaseAug and Avocodo. Different from TTS models which generate short pronunciation from phonemes and speaker identity, the category-to-sound problem requires generating diverse sounds just from a category index. To compensate for the difference while maintaining consistency within each audio clip, we heavily modified the prior encoder to enhance consistency with posterior latent variables. This introduced additional Gaussian on the prior encoder which promotes variance within the category. With these modifications, we propose VIFS, variational inference for end-to-end Foley sound synthesis, which generates diverse high-quality sounds.

📄 PDF Abstract BibTeX arXiv:2306.05004

Code (1)

junjun3518/vifs 공식 구현 pytorch

Tasks

Speech Synthesistext-to-speechText to SpeechVariational Inference

Methods 이 논문이 사용한 방법론

Variational Inference 설명 없음

Similar Papers 제목 키워드 기반

T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis

2024-01-17 · Yoonjin Chung, Junwon Lee, Juhan Nam

Foley sound, audio content inserted synchronously with videos, plays a critical role in the user experience of multimedia content. Recently, there has been active research in Foley sound synthesis, leveraging the advance…

Audio GenerationAudio Synthesis

A Proposal for Foley Sound Synthesis Challenge

2022-07-21 · Keunwoo Choi, Sangshin Oh, Minsung Kang, Brian McFee

"Foley" refers to sound effects that are added to multimedia during post-production to enhance its perceived acoustic properties, e.g., by simulating the sounds of footsteps, ambient environmental sounds, or visible obje…

Latent CLAP Loss for Better Foley Sound Synthesis

2024-03-18 · Tornike Karchkhadze, Hassan Salami Kavaki, Mohammad Rasool Izadi, Bryce Irvin 외

Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimodal text-audio representation models, su…

FAD

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

2026-03-16 · Pengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen 외 arxiv

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data,…

AutoFoley: Artificial Synthesis of Synchronized Sound Tracks for Silent Videos with Deep Learning

2020-02-21 · Sanchita Ghose, John J. Prevost

In movie productions, the Foley Artist is responsible for creating an overlay soundtrack that helps the movie come alive for the audience. This requires the artist to first identify the sounds that will enhance the exper…