paper-with-me

Papers Sound Event Detection

“Sound Event Detection” 태그가 달린 논문 211편 · 필터 해제

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

2026-07-05 · Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang 외 arxiv

Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At t…

Reinforcement LearningSound Event Detection

Semi-Supervised Sound Event Detection with Conditional Mixup and Embedding-Level Contrastive Loss

2026-06-29 · Nian Shao, Xian Li, Xiaofei Li arxiv

Sound event detection (SED) is a core module for acoustic environmental analysis, yet its performance is often limited by scarce labeled data. Recent systems leverage large pretrained audio foundation models, but effecti…

Sound Event DetectionContrastive Learning

An Analysis of Untrained Deep Reservoir Networks for Audio Surveillance

2026-06-20 · Corrado Baccheschi, Patrizio Dazzi arxiv

In this paper, we investigate untrained recurrent models from the Reservoir Computing (RC) paradigm for audio surveillance, focusing on bidirectional Echo State Networks with different depths, from shallow to deep config…

Computational EfficiencySound Event Detection

A Neuromorphic Trigger for Efficient Audio Event Detection

2026-06-16 · Benjamin Hatton, Oliver Rhodes, Luca Peres arxiv

Efficient processing of continuous audio streams remains a key challenge for real-time and resource-constrained systems. This paper introduces a neuromorphic trigger for audio event detection, based on a spiking neural n…

Sound Event Detection

Towards Open World Sound Event Detection

2026-05-05 · P. H. Hai, L. T. Minh, L. H. Son arxiv

Sound Event Detection (SED) plays a vital role in audio understanding, with applications in surveillance, smart cities, healthcare, and multimedia indexing. However, conventional SED systems operate under a closed-world …

Sound Event Detection

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video

2026-05-01 · Kazuya Tateishi, Akira Takahashi, Atsuo Hiroe, Hirofumi Takeda 외 arxiv

Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the generated audio but also explicit sound event …

Sound Event Detectionmultimodal generationAudio Generation

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt

2026-04-15 · Yanfeng Shi, Pengfei Cai, Jun Liu, Qing Gu 외 arxiv

Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferrin…

Reinforcement LearningSound Event DetectionAudio captioning

Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning

2025-12-22 · Apoorv Vyas, Heng-Jui Chang, Cheng-Fu Yang, Po-Yao Huang 외 arxiv

We introduce Perception Encoder Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Built on PE, PE-AV makes several key contributions to extend repres…

Sound Event DetectionContrastive Learning

WhaleVAD-BPN: Improving Baleen Whale Call Detection with Boundary Proposal Networks and Post-processing Optimisation

2025-10-24 · Christiaan M. Geldenhuys, Günther Tonitz, Thomas R. Niesler arxiv

While recent sound event detection (SED) systems can identify baleen whale calls in marine audio, challenges related to false positive and minority-class detection persist. We propose the boundary proposal network (BPN),…

Sound Event DetectionObject Detection

Sparse Autoencoders Make Audio Foundation Models more Explainable

2025-09-29 · Théo Mariotte, Martin Lebourdais, Antonio Almudévar, Marie Tahon 외 arxiv

Audio pretrained models are widely employed to solve various tasks in speech processing, sound event detection, or music information retrieval. However, the representations learned by these models are unclear, and their …

Self-Supervised LearningInformation RetrievalSound Event Detection

SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering

2025-09-23 · Jiarui Hai, Mounya Elhilali arxiv

Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they rem…

Sound Event DetectionData Augmentation

FlexSED: Towards Open-Vocabulary Sound Event Detection

2025-09-23 · Jiarui Hai, Helin Wang, Weizhe Guo, Mounya Elhilali arxiv

Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fundamentally limited. They cannot process f…

Multi-class ClassificationSound Event Detection

EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection

2025-09-01 · Yun Chu, Qiuhao Wang, Enze Zhou, Qian Liu 외 arxiv

Auscultation is a key method for early diagnosis of respiratory and pulmonary diseases, relying on skilled healthcare professionals. However, the process is often subjective, with variability between experts. As a result…

Sound Event DetectionGraph Neural Network

Auditory Intelligence: Understanding the World Through Sound

2025-08-11 · Hyeonuk Nam arxiv

Recent progress in auditory intelligence has yielded high-performing systems for sound event detection (SED), acoustic scene classification (ASC), automated audio captioning (AAC), and audio question answering (AQA). Yet…

Acoustic Scene ClassificationSound Event DetectionQuestion AnsweringAudio captioning

On Temporal Guidance and Iterative Refinement in Audio Source Separation

2025-07-23 · Tobias Morocutti, Jonathan Greif, Paul Primus, Florian Schmid 외 arxiv

Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a t…

Audio Source SeparationSemantic SegmentationSound Event DetectionAudio Tagging

Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries

2025-07-22 · Pengfei Cai, Yan Song, Qing Gu, Nan Jiang 외 arxiv

Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED…

Sound Event Detection

Real-Time Emergency Vehicle Siren Detection with Efficient CNNs on Embedded Hardware

2025-07-02 · Marco Giordano, Stefano Giacomelli, Claudia Rinaldi, Fabio Graziosi arxiv

We present a full-stack emergency vehicle (EV) siren detection system designed for real-time deployment on embedded hardware. The proposed approach is based on E2PANNs, a fine-tuned convolutional neural network derived f…

Sound Event Detection

Frequency Dynamic Convolutions for Sound Event Detection

2025-06-15 · Hyeonuk Nam

Recent research in deep learning-based Sound Event Detection (SED) has primarily focused on Convolutional Recurrent Neural Networks (CRNNs) and Transformer models. However, conventional 2D convolution-based models assume…

ARCEvent DetectionSound Event Detection

Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection

2025-05-27 · Shiqi Zhang, Tuomas Virtanen

Bioacoustic sound event detection (BioSED) is crucial for biodiversity conservation but faces practical challenges during model development and training: limited amounts of annotated data, sparse events, species diversit…

Active LearningDiversityEvent DetectionSound Event Detection

Exploring the Potential of SSL Models for Sound Event Detection

2025-05-17 · Hanfang Cui, Longfei Song, Li Li, Dongxing Xu 외

Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models…

Event DetectionModel SelectionSelf-Supervised LearningSound Event Detection
1–20 / 211 다음 →