paper-with-me

Papers

Label-efficient audio classification through multitask learning and self-supervision

2019-10-19 · ICLR Workshop LLD 2019 · Tyler Lee, Ting Gong, Suchismita Padhy, Andrew Rouditchenko, Anthony Ndirango

While deep learning has been incredibly successful in modeling tasks with large, carefully curated labeled datasets, its application to problems with limited labeled data remains a challenge. The aim of the present work is to improve the label efficiency of large neural networks operating on audio data through a combination of multitask learning and self-supervised learning on unlabeled data. We trained an end-to-end audio feature extractor based on WaveNet that feeds into simple, yet versatile task-specific neural networks. We describe several easily implemented self-supervised learning tasks that can operate on any large, unlabeled audio corpus. We demonstrate that, in scenarios with limited labeled training data, one can significantly improve the performance of three different supervised classification tasks individually by up to 6% through simultaneous training with these additional self-supervised tasks. We also show that incorporating data augmentation into our multitask setting leads to even further gains in performance.

📄 PDF Abstract BibTeX arXiv:1910.12587

Code (0)

등록된 구현이 없습니다.

Tasks

Audio ClassificationClassificationData AugmentationGeneral ClassificationSelf-Supervised Learning

Methods 이 논문이 사용한 방법론

Mixture of Logistic Distributions 설명 없음
Dilated Causal Convolution A Dilated Causal Convolution is a causal convolution where the filter is applied over an area larger than its length by…
WaveNet WaveNet is an audio generative model based on the PixelCNN architecture. In order to deal with long-range temporal dependencies…

Similar Papers 제목 키워드 기반

Improving label efficiency through multi-task learning on auditory data

2018-10-22 · Anonymous

Collecting high-quality, large scale datasets typically requires significant resources. The aim of the present work is to improve the label efficiency of large neural networks operating on audio data through multitask le…

Data AugmentationMulti-Task LearningSelf-Supervised Learning

Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

2026-06-11 · Olga Isupova, Danil Kuzin, Ella Browning, Tom Mills 외 arxiv

Passive acoustic monitoring holds great promise for ecological inference, yet existing automated tools are typically narrowly trained and non-transferable. We address these limitations with PULSE, a semi-supervised, mult…

Self-Supervised LearningKnowledge DistillationActive Learning

M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment

2024-03-14 · Long Nguyen-Phuoc, Renald Gaboriau, Dimitri Delacroix, Laurent Navarro

This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway ar…

On Multitask Loss Function for Audio Event Detection and Localization

2020-09-11 · Huy Phan, Lam Pham, Philipp Koch, Ngoc Q. K. Duong 외

Audio event localization and detection (SELD) have been commonly tackled using multitask models. Such a model usually consists of a multi-label event classification branch with sigmoid cross-entropy loss for event activi…

Action DetectionActivity DetectionDirection of Arrival EstimationEvent Detection+1

Dual Normalization Multitasking for Audio-Visual Sounding Object Localization

2021-06-01 · Tokuhiro Nishikawa, Daiki Shimada, Jerry Jun Yokono

Although several research works have been reported on audio-visual sound source localization in unconstrained videos, no datasets and metrics have been proposed in the literature to quantitatively evaluate its performanc…

ObjectObject LocalizationSound Source Localization