paper-with-me

Papers

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

2024-09-18 · Shota Nakada, Taichi Nishimura, Hokuto Munakata, Masayoshi Kondo, Tatsuya Komatsu

Current audio-visual representation learning can capture rough object categories (e.g., `animals'' and instruments''), but it lacks the ability to recognize fine-grained details, such as specific categories like dogs'' and `flutes'' within animals and instruments. To address this issue, we introduce DETECLAP, a method to enhance audio-visual representation learning with object information. Our key idea is to introduce an audio-visual label prediction loss to the existing Contrastive Audio-Visual Masked AutoEncoder to enhance its object awareness. To avoid costly manual annotations, we prepare object labels from both audio and visual inputs using state-of-the-art language-audio models and object detectors. We evaluate the method of audio-visual retrieval and classification using the VGGSound and AudioSet20K datasets. Our method achieves improvements in recall@10 of +1.5% and +1.2% for audio-to-visual and visual-to-audio retrieval, respectively, and an improvement in accuracy of +0.6% for audio-visual classification.

📄 PDF Abstract BibTeX arXiv:2409.11729

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectRepresentation LearningRetrieval

Similar Papers 제목 키워드 기반

Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues

2024-02-04 · Tianxiang Chen, Zhentao Tan, Tao Gong, Qi Chu 외

How to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual segmentation (AVS) task has been proposed, aiming to segment the s…

DecoderRepresentation Learning

SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering

2024-11-07 · Tianyu Yang, Yiyang Nan, Lisen Dai, Zhenwen Liang 외

Audio-Visual Question Answering (AVQA) is a challenging task that involves answering questions based on both auditory and visual information in videos. A significant challenge is interpreting complex multi-modal scenes, …

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringVisual Question Answering

Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding

2025-07-01 · Duc Cao-Dinh, Khai Le-Duc, Anh Dao, Bach Phan Tat 외 arxiv

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D…

Multi-Label ClassificationRepresentation LearningSpeech RecognitionVisual Grounding

Selective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual Segmentation

2026-03-15 · Kai Peng, Yunzhe Shen, Miao Zhang, Leiye Liu 외 arxiv

The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio-Visual Segmentation (AVS) tasks. While significant progress has been made in this area, the interaction…

Enhancing Sound Source Localization via False Negative Elimination

2024-08-29 · Zengjie Song, Jiangshe Zhang, Yuxi Wang, Junsong Fan 외

Sound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling …

audio-visual learningContrastive Learningobject-detectionObject Detection+1