DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
Current audio-visual representation learning can capture rough object categories (e.g., `animals'' and instruments''), but it lacks the ability to recognize fine-grained details, such as specific categories like dogs'' and `flutes'' within animals and instruments. To address this issue, we introduce DETECLAP, a method to enhance audio-visual representation learning with object information. Our key idea is to introduce an audio-visual label prediction loss to the existing Contrastive Audio-Visual Masked AutoEncoder to enhance its object awareness. To avoid costly manual annotations, we prepare object labels from both audio and visual inputs using state-of-the-art language-audio models and object detectors. We evaluate the method of audio-visual retrieval and classification using the VGGSound and AudioSet20K datasets. Our method achieves improvements in recall@10 of +1.5% and +1.2% for audio-to-visual and visual-to-audio retrieval, respectively, and an improvement in accuracy of +0.6% for audio-visual classification.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectRepresentation LearningRetrievalSimilar Papers 제목 키워드 기반
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
How to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual segmentation (AVS) task has been proposed, aiming to segment the s…
DecoderRepresentation LearningSaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
Audio-Visual Question Answering (AVQA) is a challenging task that involves answering questions based on both auditory and visual information in videos. A significant challenge is interpreting complex multi-modal scenes, …
Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringVisual Question AnsweringAudio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D…
Multi-Label ClassificationRepresentation LearningSpeech RecognitionVisual GroundingSelective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual Segmentation
The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio-Visual Segmentation (AVS) tasks. While significant progress has been made in this area, the interaction…
Enhancing Sound Source Localization via False Negative Elimination
Sound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling …
audio-visual learningContrastive Learningobject-detectionObject Detection+1