paper-with-me

Papers

wav2pos: Sound Source Localization using Masked Autoencoders

2024-08-28 · Axel Berg, Jens Gulin, Mark O'Connor, Chuteng Zhou, Karl Åström, Magnus Oskarsson

We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods.

📄 PDF Abstract BibTeX arXiv:2408.15771

Code (1)

axeber01/wav2pos 공식 구현 pytorch

Tasks

Indoor LocalizationSound Source Localization

Similar Papers 제목 키워드 기반

Masked Autoencoders for Egocentric Video Understanding @ Ego4D Challenge 2022

2022-11-18 · Jiachen Lei, Shuang Ma, Zhongjie Ba, Sai Vemprala 외

In this report, we present our approach and empirical results of applying masked autoencoders in two egocentric video understanding tasks, namely, Object State Change Classification and PNR Temporal Localization, of Ego4…

Object State Change ClassificationTemporal LocalizationVideo Understanding

AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers

2025-06-03 · Linya Fu, Yu Liu, Zhijie Liu, Zedong Yang 외

We propose AuralNet, a novel 3D multi-source binaural sound source localization approach that localizes overlapping sources in both azimuth and elevation without prior knowledge of the number of sources. AuralNet employs…

Sound Source Localization

Focus on Texture: Rethinking Pre-training in Masked Autoencoders for Medical Image Classification

2025-07-15 · Chetan Madan, Aarjav Satia, Soumen Basu, Pankaj Gupta 외 arxiv

Masked Autoencoders (MAEs) have emerged as a dominant strategy for self-supervised representation learning in natural images, where models are pre-trained to reconstruct masked patches with a pixel-wise mean squared erro…

Gallbladder Cancer DetectionMedical Image ClassificationUnsupervised Pre-trainingRepresentation Learning

Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications

2025-08-28 · Immanuel Roßteutscher, Klaus S. Drese, Thorsten Uphues arxiv

We investigated the adaptation and performance of Masked Autoencoders (MAEs) with Vision Transformer (ViT) architectures for self-supervised representation learning on one-dimensional (1D) ultrasound signals. Although MA…

Self-Supervised LearningRepresentation Learning

Sound Source Localization is All about Cross-Modal Alignment

2023-09-19 · ICCV 2023 1 · Arda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 외

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localizati…

Allcross-modal alignmentCross-Modal RetrievalRetrieval+1