Text-Queried Target Sound Event Localization
Sound event localization and detection (SELD) aims to determine the appearance of sound classes, together with their Direction of Arrival (DOA). However, current SELD systems can only predict the activities of specific classes, for example, 13 classes in DCASE challenges. In this paper, we propose text-queried target sound event localization (SEL), a new paradigm that allows the user to input the text to describe the sound event, and the SEL model can predict the location of the related sound event. The proposed task presents a more user-friendly way for human-computer interaction. We provide a benchmark study for the proposed task and perform experiments on datasets created by simulated room impulse response (RIR) and real RIR to validate the effectiveness of the proposed methods. We hope that our benchmark will inspire the interest and additional research for text-queried sound source localization.
Code (0)
등록된 구현이 없습니다.
Tasks
Room Impulse Response (RIR)Sound Event Localization and DetectionSound Source LocalizationSimilar Papers 제목 키워드 기반
Leveraging LLM and Text-Queried Separation for Noise-Robust Sound Event Detection
Sound Event Detection (SED) is challenging in noisy environments where overlapping sounds obscure target events. Language-queried audio source separation (LASS) aims to isolate the target sound events from a noisy clip. …
Audio Source SeparationEvent DetectionSound Event DetectionExploring Text-Queried Sound Event Detection with Audio Source Separation
In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this …
Audio Source SeparationEvent DetectionSound Event DetectionCLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos
Recent years have seen progress beyond domain-specific sound separation for speech or music towards universal sound separation for arbitrary sounds. Prior work on universal sound separation has investigated separating a …
FPAN: Fine-grained and Progressive Attention Localization Network for Data Retrieval
The Localization of the target object for data retrieval is a key issue in the Intelligent and Connected Transportation Systems (ICTS). However, due to lack of intelligence in the traditional transportation system, it ca…
Multi-Task LearningObjectObject LocalizationObject Tracking+1Separate Anything You Describe
Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provide…
Audio Source SeparationNatural Language QueriesSpeech EnhancementZero-shot Generalization