paper-with-me

Papers

Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification

2025-07-16 · Kazuki Shimada, Archontis Politis, Iran R. Roman, Parthasaarathy Sudarsanam, David Diaz-Guerra, Ruchi Pandey, Kengo Uchida, Yuichiro Koyama, Naoya Takahashi, Takashi Shibuya, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji arxiv

This paper presents the objective, dataset, baseline, and metrics of Task 3 of the DCASE2025 Challenge on sound event localization and detection (SELD). In previous editions, the challenge used four-channel audio formats of first-order Ambisonics (FOA) and microphone array. In contrast, this year's challenge investigates SELD with stereo audio data (termed stereo SELD). This change shifts the focus from more specialized 360° audio and audiovisual scene analysis to more commonplace audio and media scenarios with limited field-of-view (FOV). Due to inherent angular ambiguities in stereo audio data, the task focuses on direction-of-arrival (DOA) estimation in the azimuth plane (left-right axis) along with distance estimation. The challenge remains divided into two tracks: audio-only and audiovisual, with the audiovisual track introducing a new sub-task of onscreen/offscreen event classification necessitated by the limited FOV. This challenge introduces the DCASE2025 Task3 Stereo SELD Dataset, whose stereo audio and perspective video clips are sampled and converted from the STARSS23 recordings. The baseline system is designed to process stereo audio and corresponding video frames as inputs. In addition to the typical SELD event classification and localization, it integrates onscreen/offscreen classification for the audiovisual track. The evaluation metrics have been modified to introduce an onscreen/offscreen accuracy metric, which assesses the models' ability to identify which sound sources are onscreen. In the experimental evaluation, the baseline system performs reasonably well with the stereo audio data.

📄 PDF Abstract BibTeX arXiv:2507.12042

Code (0)

등록된 구현이 없습니다.

Tasks

Sound Event Localization and Detection

Similar Papers 제목 키워드 기반

Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling

2025-06-16 · Wenmiao Gao, Yang Xiao

Pre-training methods have achieved significant performance improvements in sound event localization and detection (SELD) tasks, but existing Transformer-based models suffer from high computational complexity. In this wor…

DecoderMambaSound Event Localization and Detection

ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video

2026-01-24 · Davide Berghi, Philip J. B. Jackson arxiv

Sound event localization and detection with distance estimation (3D SELD) in video involves identifying active sound events at each time frame while estimating their spatial coordinates. This multimodal task requires joi…

Sound Event Localization and Detection

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

2024-12-14 · CVPR 2025 1 · Saksham Singh Kushwaha, Yapeng Tian

Recent advances in audio generation have focused on text-to-audio (T2A) and video-to-audio (V2A) tasks. However, T2A or V2A methods cannot generate holistic sounds (onscreen and off-screen). This is because T2A cannot ge…

Audio Generation

Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos

2025-09-08 · Davide Berghi, Philip J. B. Jackson arxiv

In this study, we address the multimodal task of stereo sound event localization and detection with source distance estimation (3D SELD) in regular video content. 3D SELD is a complex task that combines temporal event cl…

Sound Event Localization and Detection

Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos

2025-07-07 · Davide Berghi, Philip J. B. Jackson

This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. SELD is a complex tas…

Sound Event Localization and Detection