paper-with-me

Papers

Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation

2024-06-18 · Miseul Kim, Soo-Whan Chung, Youna Ji, Hong-Goo Kang, Min-Seok Choi

This paper introduces a novel task in generative speech processing, Acoustic Scene Transfer (AST), which aims to transfer acoustic scenes of speech signals to diverse environments. AST promises an immersive experience in speech perception by adapting the acoustic scene behind speech signals to desired environments. We propose AST-LDM for the AST task, which generates speech signals accompanied by the target acoustic scene of the reference prompt. Specifically, AST-LDM is a latent diffusion model conditioned by CLAP embeddings that describe target acoustic scenes in either audio or text modalities. The contributions of this paper include introducing the AST task and implementing its baseline model. For AST-LDM, we emphasize its core framework, which is to preserve the input speech and generate audio consistently with both the given speech and the target acoustic environment. Experiments, including objective and subjective tests, validate the feasibility and efficacy of our approach.

📄 PDF Abstract BibTeX arXiv:2406.12688

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Latent Diffusion Model Diffusion models applied to latent spaces, which are normally built with (Variational) Autoencoders.

Similar Papers 제목 키워드 기반

Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments

2021-06-30 · Stefan Fichna, Thomas Biberger, Bernhard U. Seeber, Stephan D. Ewert

In daily life, social interaction and acoustic communication often take place in complex acoustic environments (CAE) with a variety of interfering sounds and reverberation. For hearing research and the evaluation of hear…

Characterizing dynamically varying acoustic scenes from egocentric audio recordings in workplace setting

2019-11-10 · Arindam Jati, Amrutha Nadarajan, Karel Mundnich, Shrikanth Narayanan

Devices capable of detecting and categorizing acoustic scenes have numerous applications such as providing context-aware user experiences. In this paper, we address the task of characterizing acoustic scenes in a workpla…

Acoustic Scene ClassificationGeneral ClassificationScene Classification

Constrained speaker diarization of TV series based on visual patterns

2018-12-18 · Xavier Bost, Georges Linares

Speaker diarization, usually denoted as the ''who spoke when'' task, turns out to be particularly challenging when applied to fictional films, where many characters talk in various acoustic conditions (background music, …

Clusteringspeaker-diarizationSpeaker Diarization

D{é}tection de locuteurs dans les s{é}ries TV

2018-12-18 · Xavier Bost, Georges Linares

Speaker diarization of audio streams turns out to be particularly challenging when applied to fictional films, where many characters talk in various acoustic conditions (background music, sound effects, variations in int…

Clusteringspeaker-diarizationSpeaker Diarization

Acoustic Field Video for Multimodal Scene Understanding

2026-01-23 · Daehwa Kim, Chris Harrison arxiv

We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video stream provides a spatially grounded visu…

Multimodal ReasoningScene Understanding