paper-with-me

홈 › Papers

Sound Localization by Self-Supervised Time Delay Estimation

2022-04-26 · Ziyang Chen, David F. Fouhey, Andrew Owens

Sounds reach one microphone in a stereo pair sooner than the other, resulting in an interaural time delay that conveys their directions. Estimating a sound's time delay requires finding correspondences between the signals recorded by each microphone. We propose to learn these correspondences through self-supervision, drawing on recent techniques from visual tracking. We adapt the contrastive random walk of Jabri et al. to learn a cycle-consistent representation from unlabeled stereo sounds, resulting in a model that performs on par with supervised methods on "in the wild" internet recordings. We also propose a multimodal contrastive learning model that solves a visually-guided localization task: estimating the time delay for a particular person in a multi-speaker mixture, given a visual representation of their face. Project site: https://ificl.github.io/stereocrw/

📄 PDF Abstract BibTeX arXiv:2204.12489

Code (1)

IFICL/stereocrw 공식 구현 pytorch

Tasks

Contrastive LearningVisual Tracking

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

2020-10-12 · NeurIPS 2020 12 · Di Hu, Rui Qian, Minyue Jiang, Xiao Tan 외

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform…

ObjectObject Localization

Hear The Flow: Optical Flow-Based Self-Supervised Visual Sound Source Localization

2022-11-06 · Dennis Fedorishin, Deen Dayal Mohan, Bhavin Jawade, Srirangaraj Setlur 외

Learning to localize the sound source in videos without explicit annotations is a novel area of audio-visual research. Existing work in this area focuses on creating attention maps to capture the correlation between the …

Optical Flow EstimationSound Source Localization

Self-Supervised Predictive Learning: A Negative-Free Method for Sound Source Localization in Visual Scenes

2022-03-25 · CVPR 2022 1 · Zengjie Song, Yuxi Wang, Junsong Fan, Tieniu Tan 외

Sound source localization in visual scenes aims to localize objects emitting the sound in a given image. Recent works showing impressive localization performance typically rely on the contrastive learning framework. Howe…

Contrastive LearningSound Source Localization

Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation

2023-03-20 · ICCV 2023 1 · Ziyang Chen, Shengyi Qian, Andrew Owens

The images and sounds that we perceive undergo subtle but geometrically consistent changes as we rotate our heads. In this paper, we use these cues to solve a problem we call Sound Localization from Motion (SLfM): jointl…

BeamLearning: an end-to-end Deep Learning approach for the angular localization of sound sources using raw multichannel acoustic pressure data

2021-04-27 · Hadrien Pujol, Éric Bavu, Alexandre Garcia

Sound sources localization using multichannel signal processing has been a subject of active research for decades. In recent years, the use of deep learning in audio signal processing has allowed to drastically improve p…

Audio Signal ProcessingBIG-bench Machine LearningComputational EfficiencyGPU