paper-with-me

Papers

Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization

2023-08-09 · Tianyu Liu, Peng Zhang, Wei Huang, Yufei zha, Tao You, Yanning Zhang

Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to establish such a consistent correspondence between audio and sound sources in visual scenarios. Unfortunately, the insufficient attention to the heterogeneity influence in the different modality features still limits this scheme to be further improved, which also becomes the motivation of our work. In this study, an Induction Network is proposed to bridge the modality gap more effectively. By decoupling the gradients of visual and audio modalities, the discriminative visual representations of sound sources can be learned with the designed Induction Vector in a bootstrap manner, which also enables the audio modality to be aligned with the visual modality consistently. In addition to a visual weighted contrastive loss, an adaptive threshold selection strategy is introduced to enhance the robustness of the Induction Network. Substantial experiments conducted on SoundNet-Flickr and VGG-Sound Source datasets have demonstrated a superior performance compared to other state-of-the-art works in different challenging scenarios. The code is available at https://github.com/Tahy1/AVIN

📄 PDF Abstract BibTeX arXiv:2308.04767

Code (1)

tahy1/avin 공식 구현 pytorch

Tasks

Contrastive LearningSound Source Localization

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video

2022-04-04 · ICCV 2021 10 · Minsu Kim, Joanna Hong, Se Jin Park, Yong Man Ro

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual…

Lip Reading

Grammar Induction from Visual, Speech and Text

2024-10-01 · Yu Zhao, Hao Fei, Shengqiong Wu, Meishan Zhang 외

Grammar Induction could benefit from rich heterogeneous signals, such as text, vision, and acoustics. In the process, features from distinct modalities essentially serve complementary roles to each other. With such intui…

Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement

2025-01-23 · Meng-Ping Lin, Jen-Cheng Hou, Chia-Wei Chen, Shao-Yi Chien 외

Speech enhancement (SE) aims to improve the quality and intelligibility of speech in noisy environments. Recent studies have shown that incorporating visual cues in audio signal processing can enhance SE performance. Giv…

Audio Signal ProcessingSpeech EnhancementTransfer Learning

AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model

2023-08-15 · Jeong Hun Yeo, Minsu Kim, Jeongsoo Choi, Dae Hoe Kim 외

Visual Speech Recognition (VSR) is the task of predicting spoken words from silent lip movements. VSR is regarded as a challenging task because of the insufficient information on lip movements. In this paper, we propose …

Quantizationspeech-recognitionSpeech RecognitionVisual Speech Recognition

Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

2025-10-13 · KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo 외 arxiv

Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Dif…

Audio captioning