paper-with-me

Papers

Few-Shot Audio-Visual Learning of Environment Acoustics

2022-06-08 · Sagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen Grauman

Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate RIRs assume dense geometry and/or sound measurements throughout the environment, we explore how to infer RIRs based on a sparse set of images and echoes observed in the space. Towards that goal, we introduce a transformer-based method that uses self-attention to build a rich acoustic context, then predicts RIRs of arbitrary query source-receiver locations through cross-attention. Additionally, we design a novel training objective that improves the match in the acoustic signature between the RIR predictions and the targets. In experiments using a state-of-the-art audio-visual simulator for 3D environments, we demonstrate that our method successfully generates arbitrary RIRs, outperforming state-of-the-art methods and -- in a major departure from traditional methods -- generalizing to novel environments in a few-shot manner. Project: http://vision.cs.utexas.edu/projects/fs_rir.

📄 PDF Abstract BibTeX arXiv:2206.04006

Code (0)

등록된 구현이 없습니다.

Tasks

audio-visual learningRoom Impulse Response (RIR)

Similar Papers 제목 키워드 기반

MAGIC: Map-Guided Few-Shot Audio-Visual Acoustics Modeling

2024-05-22 · Diwei Huang, Kunyang Lin, Peihao Chen, Qing Du 외

Few-shot audio-visual acoustics modeling seeks to synthesize the room impulse response in arbitrary locations with few-shot observations. To sufficiently exploit the provided few-shot data for accurate acoustic modeling,…

Decoder

NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics

2024-11-11 · David Robinson, Marius Miron, Masato Hagiwara, Olivier Pietquin

Large language models (LLMs) prompted with text and audio represent the state of the art in various auditory tasks, including speech, music, and general audio, showing emergent abilities on unseen tasks. However, these c…

zero-shot-classificationZero-Shot Learning

Visual Acoustic Matching

2022-02-14 · CVPR 2022 1 · Changan Chen, Ruohan Gao, Paul Calamia, Kristen Grauman

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, t…

Self-Supervised Visual Acoustic Matching

2023-07-27 · NeurIPS 2023 11

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source a…

Diversity

Self-Supervised Learning for Few-Shot Bird Sound Classification

2023-12-25 · Ilyass Moummad, Romain Serizel, Nicolas Farrugia

Self-supervised learning (SSL) in audio holds significant potential across various domains, particularly in situations where abundant, unlabeled data is readily available at no cost. This is pertinent in bioacoustics, wh…

ClassificationFew-Shot LearningSelf-Supervised LearningSound Classification