Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
For audio in augmented reality (AR), knowledge of the users' real acoustic environment is crucial for rendering virtual sounds that seamlessly blend into the environment. As acoustic measurements are usually not feasible in practical AR applications, information about the room needs to be inferred from available sound sources. Then, additional sound sources can be rendered with the same room acoustic qualities. Crucially, these are placed at different positions than the sources available for estimation. Here, we propose to use an encoder network trained using a contrastive loss that maps input sounds to a low-dimensional feature space representing only room-specific information. Then, a diffusion-based spatial room impulse response generator is trained to take the latent space and generate a new response, given a new source-receiver position. We show how both room- and position-specific parameters are considered in the final output.
Code (0)
등록된 구현이 없습니다.
Tasks
PositionResponse GenerationSimilar Papers 제목 키워드 기반
Blind identification of Ambisonic reduced room impulse response
Recently proposed Generalized Time-domain Velocity Vector (GTVV) is a generalization of relative room impulse response in spherical harmonic (aka Ambisonic) domain that allows for blind estimation of early-echo parameter…
Room Impulse Response (RIR)Time SeriesBlind Localization of Room Reflections with Application to Spatial Audio
Blind estimation of early room reflections, without knowledge of the room impulse response, holds substantial value. The FF-PHALCOR (Frequency Focusing PHase ALigned CORrelation), method was recently developed for this o…
BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion Models
In this paper, we present an unsupervised single-channel method for joint blind dereverberation and room impulse response estimation, based on posterior sampling with diffusion models. We parameterize the reverberation o…
Blind Localization of Early Room Reflections with Arbitrary Microphone Array
Blindly estimating the direction of arrival (DoA) of early room reflections without prior knowledge of the room impulse response or source signal is highly valuable in audio signal processing applications. The FF-PHALCOR…
Audio Signal ProcessingDiffusion Posterior Sampling for Informed Single-Channel Dereverberation
We present in this paper an informed single-channel dereverberation method based on conditional generation with diffusion models. With knowledge of the room impulse response, the anechoic utterance is generated via rever…