paper-with-me

Papers

DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection

2024-10-30 · Yoto Fujita, Yoshiaki Bando, Keisuke Imoto, Masaki Onishi, Kazuyoshi Yoshii

This paper describes sound event localization and detection (SELD) for spatial audio recordings captured by firstorder ambisonics (FOA) microphones. In this task, one may train a deep neural network (DNN) using FOA data annotated with the classes and directions of arrival (DOAs) of sound events. However, the performance of this approach is severely bounded by the amount of annotated data. To overcome this limitation, we propose a novel method of pretraining the feature extraction part of the DNN in a self-supervised manner. We use spatial audio-visual recordings abundantly available as virtual reality contents. Assuming that sound objects are concurrently observed by the FOA microphones and the omni-directional camera, we jointly train audio and visual encoders with contrastive learning such that the audio and visual embeddings of the same recording and DOA are made close. A key feature of our method is that the DOA-wise audio embeddings are jointly extracted from the raw audio data, while the DOA-wise visual embeddings are separately extracted from the local visual crops centered on the corresponding DOA. This encourages the latent features of the audio encoder to represent both the classes and DOAs of sound events. The experiment using the DCASE2022 Task 3 dataset of 20 hours shows non-annotated audio-visual recordings of 100 hours reduced the error score of SELD from 36.4 pts to 34.9 pts.

📄 PDF Abstract BibTeX arXiv:2410.22803

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningSelf-Supervised LearningSound Event Localization and Detection

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

2020-10-12 · NeurIPS 2020 12 · Di Hu, Rui Qian, Minyue Jiang, Xiao Tan 외

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform…

ObjectObject Localization

Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling

2020-07-28 · Yoshiki Masuyama, Yoshiaki Bando, Kohei Yatabe, Yoko Sasaki 외

Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living env…

Self-Supervised LearningSound Source Localization

Self-supervised Contrastive Learning for Audio-Visual Action Recognition

2022-04-28 · Yang Liu, Ying Tan, Haoyuan Lan

The underlying correlation between audio and visual modalities can be utilized to learn supervised information for unlabeled videos. In this paper, we propose an end-to-end self-supervised framework named Audio-Visual Co…

Action RecognitionContrastive LearningSelf-Supervised Action Recognition

Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization

2025-05-08 · Sooyoung Park, Arda Senocak, Joon Son Chung

Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the applic…

Scene UnderstandingSound Source Localization

Telling Left from Right: Learning Spatial Correspondence of Sight and Sound

2020-06-11 · CVPR 2020 6 · Karren Yang, Bryan Russell, Justin Salamon

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic informa…

audio-visual learning