Papers Sound Event Detection
“Sound Event Detection” 태그가 달린 논문 211편 · 필터 해제
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At t…
Reinforcement LearningSound Event DetectionSemi-Supervised Sound Event Detection with Conditional Mixup and Embedding-Level Contrastive Loss
Sound event detection (SED) is a core module for acoustic environmental analysis, yet its performance is often limited by scarce labeled data. Recent systems leverage large pretrained audio foundation models, but effecti…
Sound Event DetectionContrastive LearningAn Analysis of Untrained Deep Reservoir Networks for Audio Surveillance
In this paper, we investigate untrained recurrent models from the Reservoir Computing (RC) paradigm for audio surveillance, focusing on bidirectional Echo State Networks with different depths, from shallow to deep config…
Computational EfficiencySound Event DetectionA Neuromorphic Trigger for Efficient Audio Event Detection
Efficient processing of continuous audio streams remains a key challenge for real-time and resource-constrained systems. This paper introduces a neuromorphic trigger for audio event detection, based on a spiking neural n…
Sound Event DetectionTowards Open World Sound Event Detection
Sound Event Detection (SED) plays a vital role in audio understanding, with applications in surveillance, smart cities, healthcare, and multimedia indexing. However, conventional SED systems operate under a closed-world …
Sound Event DetectionMMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video
Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the generated audio but also explicit sound event …
Sound Event Detectionmultimodal generationAudio GenerationTowards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt
Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferrin…
Reinforcement LearningSound Event DetectionAudio captioningPushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
We introduce Perception Encoder Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Built on PE, PE-AV makes several key contributions to extend repres…
Sound Event DetectionContrastive LearningWhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation
While recent sound event detection (SED) systems can identify baleen whale calls in marine audio, challenges related to false positive and minority-class detection persist. We propose the boundary proposal network (BPN),…
Sound Event DetectionObject DetectionSparse Autoencoders Make Audio Foundation Models more Explainable
Audio pretrained models are widely employed to solve various tasks in speech processing, sound event detection, or music information retrieval. However, the representations learned by these models are unclear, and their …
Self-Supervised LearningInformation RetrievalSound Event DetectionSynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering
Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they rem…
Sound Event DetectionData AugmentationFlexSED: Towards Open-Vocabulary Sound Event Detection
Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fundamentally limited. They cannot process f…
Multi-class ClassificationSound Event DetectionEZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
Auscultation is a key method for early diagnosis of respiratory and pulmonary diseases, relying on skilled healthcare professionals. However, the process is often subjective, with variability between experts. As a result…
Sound Event DetectionGraph Neural NetworkAuditory Intelligence: Understanding the World Through Sound
Recent progress in auditory intelligence has yielded high-performing systems for sound event detection (SED), acoustic scene classification (ASC), automated audio captioning (AAC), and audio question answering (AQA). Yet…
Acoustic Scene ClassificationSound Event DetectionQuestion AnsweringAudio captioningOn Temporal Guidance and Iterative Refinement in Audio Source Separation
Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a t…
Audio Source SeparationSemantic SegmentationSound Event DetectionAudio TaggingDetect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED…
Sound Event DetectionReal-Time Emergency Vehicle Siren Detection with Efficient CNNs on Embedded Hardware
We present a full-stack emergency vehicle (EV) siren detection system designed for real-time deployment on embedded hardware. The proposed approach is based on E2PANNs, a fine-tuned convolutional neural network derived f…
Sound Event DetectionFrequency Dynamic Convolutions for Sound Event Detection
Recent research in deep learning-based Sound Event Detection (SED) has primarily focused on Convolutional Recurrent Neural Networks (CRNNs) and Transformer models. However, conventional 2D convolution-based models assume…
ARCEvent DetectionSound Event DetectionHybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
Bioacoustic sound event detection (BioSED) is crucial for biodiversity conservation but faces practical challenges during model development and training: limited amounts of annotated data, sparse events, species diversit…
Active LearningDiversityEvent DetectionSound Event DetectionExploring the Potential of SSL Models for Sound Event Detection
Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models…
Event DetectionModel SelectionSelf-Supervised LearningSound Event Detection