paper-with-me

홈 › Papers

Learning to Detect Novel and Fine-Grained Acoustic Sequences Using Pretrained Audio Representations

2023-05-03 · Vasudha Kowtha, Miquel Espi Marques, Jonathan Huang, Yichi Zhang, Carlos Avendano

This work investigates pretrained audio representations for few shot Sound Event Detection. We specifically address the task of few shot detection of novel acoustic sequences, or sound events with semantically meaningful temporal structure, without assuming access to non-target audio. We develop procedures for pretraining suitable representations, and methods which transfer them to our few shot learning scenario. Our experiments evaluate the general purpose utility of our pretrained representations on AudioSet, and the utility of proposed few shot methods via tasks constructed from real-world acoustic sequences. Our pretrained embeddings are suitable to the proposed task, and enable multiple aspects of our few shot framework.

📄 PDF Abstract BibTeX arXiv:2305.02382

Code (0)

등록된 구현이 없습니다.

Tasks

Event DetectionFew-Shot LearningSound Event Detection

Similar Papers 제목 키워드 기반

Audiovisual transfer learning for audio tagging and sound event detection

2021-06-09 · Wim Boes, Hugo Van hamme

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to a…

Audio TaggingEvent DetectionSound Event DetectionTransfer Learning

Acoustic models of Brazilian Portuguese Speech based on Neural Transformers

2023-12-14 · Marcelo Matheus Gauy, Marcelo Finger

An acoustic model, trained on a significant amount of unlabeled data, consists of a self-supervised learned speech representation useful for solving downstream tasks, perhaps after a fine-tuning of the model in the respe…

Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech

2025-10-05 · Yuxin Li, Eng Siong Chng, Cuntai Guan arxiv

Speech-based depression detection (SDD) has emerged as a non-invasive and scalable alternative to conventional clinical assessments. However, existing methods still struggle to capture robust depression-related speech ch…

Self-Supervised LearningRepresentation Learning

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

2026-06-03 · Zhihan Li, Hankun Wang, Yiwei Guo, Bohan Li 외 arxiv

Automatic speech recognition systems commonly rely on reference transcriptions for evaluation, while reference-free approaches often depend on internal confidence estimation or auxiliary language models. We propose READ …

Speech Recognition

Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study

2026-05-13 · Wuao Liu, Mustafa Chasmai, Subhransu Maji, Grant Van Horn arxiv

Bioacoustic recognition requires fine-grained acoustic understanding to distinguish similar-sounding species. However, many large-scale data repositories such as iNaturalist are weakly annotated, often with only a single…

Self-Supervised Learning