paper-with-me

홈 › Papers

MAGIC: Map-Guided Few-Shot Audio-Visual Acoustics Modeling

2024-05-22 · Diwei Huang, Kunyang Lin, Peihao Chen, Qing Du, Mingkui Tan

Few-shot audio-visual acoustics modeling seeks to synthesize the room impulse response in arbitrary locations with few-shot observations. To sufficiently exploit the provided few-shot data for accurate acoustic modeling, we present a *map-guided* framework by constructing acoustic-related visual semantic feature maps of the scenes. Visual features preserve semantic details related to sound and maps provide explicit structural regularities of sound propagation, which are valuable for modeling environment acoustics. We thus extract pixel-wise semantic features derived from observations and project them into a top-down map, namely the observation semantic map. This map contains the relative positional information among points and the semantic feature information associated with each point. Yet, limited information extracted by few-shot observations on the map is not sufficient for understanding and modeling the whole scene. We address the challenge by generating a scene semantic map via diffusing features and anticipating the observation semantic map. The scene semantic map then interacts with echo encoding by a transformer-based encoder-decoder to predict RIR for arbitrary speaker-listener query pairs. Extensive experiments on Matterport3D and Replica dataset verify the efficacy of our framework.

📄 PDF Abstract BibTeX arXiv:2405.13860

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control

2025-10-26 · Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia, Muhammad Awais 외 arxiv

Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consi…

Talking Face GenerationVideo Generation

Language Models Can See: Plugging Visual Controls in Text Generation

2022-05-05 · Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu 외

Generative language models (LMs) such as GPT-2/3 can be prompted to generate text with remarkable quality. While they are designed for text-prompted generation, it remains an open question how the generation process coul…

Image CaptioningImage-text matchingOpen-Ended Question AnsweringStory Generation+2

NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics

2024-11-11 · David Robinson, Marius Miron, Masato Hagiwara, Olivier Pietquin

Large language models (LLMs) prompted with text and audio represent the state of the art in various auditory tasks, including speech, music, and general audio, showing emergent abilities on unseen tasks. However, these c…

zero-shot-classificationZero-Shot Learning

MAGIC-Enhanced Keyword Prompting for Zero-Shot Audio Captioning with CLIP Models

2025-09-16 · Vijay Govindarajan, Pratik Patel, Sahil Tripathi, Md Azizul Hoque 외 arxiv

Automated Audio Captioning (AAC) generates captions for audio clips but faces challenges due to limited datasets compared to image captioning. To overcome this, we propose the zero-shot AAC system that leverages pre-trai…

Zero-shot Audio CaptioningImage Captioning

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

2025-08-04 · Xu Wang, Shengeng Tang, Fei Wang, Lechao Cheng 외 arxiv

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-qu…

Talking Face Generation