paper-with-me

홈 › Papers

Unified Audio Event Detection

2024-09-13 · Yidi Jiang, Ruijie Tao, Wen Huang, Qian Chen, Wen Wang

Sound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech event, while in SD, non-speech sounds are treated merely as background noise. Thus, both tasks provide only partial analysis in complex audio scenarios involving both speech conversation and non-speech sounds. In this paper, we introduce a novel task called Unified Audio Event Detection (UAED) for comprehensive audio analysis. UAED explores the synergy between SED and SD tasks, simultaneously detecting non-speech sound events and fine-grained speech events based on speaker identities. To tackle this task, we propose a Transformer-based UAED (T-UAED) framework and construct the UAED Data derived from the Librispeech dataset and DESED soundbank. Experiments demonstrate that the proposed framework effectively exploits task interactions and substantially outperforms the baseline that simply combines the outputs of SED and SD models. T-UAED also shows its versatility by performing comparably to specialized models for individual SED and SD tasks on DESED and CALLHOME datasets.

📄 PDF Abstract BibTeX arXiv:2409.08552

Code (0)

등록된 구현이 없습니다.

Tasks

Event DetectionSound Event Detectionspeaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization

2024-04-04 · Tiantian Geng, Teng Wang, yanfu Zhang, Jinming Duan 외

Video localization tasks aim to temporally locate specific instances in videos, including temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL). Existing methods over-…

Action Localizationaudio-visual event localizationEvent DetectionSound Event Detection+1

Real-Time Emergency Vehicle Siren Detection with Efficient CNNs on Embedded Hardware

2025-07-02 · Marco Giordano, Stefano Giacomelli, Claudia Rinaldi, Fabio Graziosi arxiv

We present a full-stack emergency vehicle (EV) siren detection system designed for real-time deployment on embedded hardware. The proposed approach is based on E2PANNs, a fine-tuned convolutional neural network derived f…

Sound Event Detection

Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs

2026-04-14 · Linhao Zhang, Yuhan Song, Aiwei Liu, Chuhan Wu 외 arxiv

Recent Audio Large Language Models (AudioLLMs) exhibit a striking performance inversion: while excelling at complex reasoning tasks, they consistently underperform on fine-grained acoustic perception. We attribute this g…

NowYouSee Me: Context-Aware Automatic Audio Description

2024-12-13 · Seon-Ho Lee, Jue Wang, David Fan, Zhikang Zhang 외

Audio Description (AD) plays a pivotal role as an application system aimed at guaranteeing accessibility in multimedia content, which provides additional narrations at suitable intervals to describe visual elements, cate…

Event DetectionScript Generation

Audio Event and Scene Recognition: A Unified Approach using Strongly and Weakly Labeled Data

2016-11-12 · Anurag Kumar, Bhiksha Raj

In this paper we propose a novel learning framework called Supervised and Weakly Supervised Learning where the goal is to learn simultaneously from weakly and strongly labeled data. Strongly labeled data can be simply un…

Scene RecognitionWeakly-supervised Learning