paper-with-me

Papers

Improved Zero-Shot Audio Tagging & Classification with Patchout Spectrogram Transformers

2022-08-24 · Paul Primus, Gerhard Widmer

Standard machine learning models for tagging and classifying acoustic signals cannot handle classes that were not seen during training. Zero-Shot (ZS) learning overcomes this restriction by predicting classes based on adaptable class descriptions. This study sets out to investigate the effectiveness of self-attention-based audio embedding architectures for ZS learning. To this end, we compare the very recent patchout spectrogram transformer with two classic convolutional architectures. We evaluate these three architectures on three tasks and on three different benchmark datasets: general-purpose tagging on AudioSet, environmental sound classification on ESC-50, and instrument tagging on OpenMIC. Our results show that the self-attention-based embedding methods outperform both compared convolutional architectures in all of these settings. By designing training and test data accordingly, we observe that prediction performance suffers significantly when the `semantic distance' between training and new test classes is large, an effect that will deserve more detailed investigations.

📄 PDF Abstract BibTeX arXiv:2208.11402

Code (0)

등록된 구현이 없습니다.

Tasks

Audio TaggingClassificationEnvironmental Sound ClassificationSound Classification

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Zero-shot Learning and Knowledge Transfer in Music Classification and Tagging

2019-06-20 · Jeong Choi, Jongpil Lee, Jiyoung Park, Juhan Nam

Music classification and tagging is conducted through categorical supervised learning with a fixed set of labels. In principle, this cannot make predictions on unseen labels. Zero-shot learning is an approach to solve th…

ClassificationGeneral ClassificationMusic ClassificationTransfer Learning+1

Zero-shot Learning for Audio-based Music Classification and Tagging

2019-07-05 · Jeong Choi, Jongpil Lee, Jiyoung Park, Juhan Nam

Audio-based music classification and tagging is typically based on categorical supervised learning with a fixed set of labels. This intrinsically cannot handle unseen labels such as newly added music genres or semantic w…

AttributeClassificationGeneral ClassificationMulti-label zero-shot learning+4

Joint Music and Language Attention Models for Zero-shot Music Tagging

2023-10-16 · Xingjian Du, Zhesong Yu, Jiaju Lin, Bilei Zhu 외

Music tagging is a task to predict the tags of music recordings. However, previous music tagging research primarily focuses on close-set music tagging tasks which can not be generalized to new tags. In this work, we prop…

Audio TaggingDecoderMusic Tagging

Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge Transfer

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Machines that can represent and describe environmental soundscapes have practical potential, e.g., for audio tagging and captioning. Prevailing learning paradigms of audio-text connections have been relying on parallel a…

Audio ClassificationAudio TaggingRetrievalTransfer Learning+1

Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge Transfer

2021-12-16 · NAACL 2022 7 · Yanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu 외

Machines that can represent and describe environmental soundscapes have practical potential, e.g., for audio tagging and captioning systems. Prevailing learning paradigms have been relying on parallel audio-text data, wh…

Audio ClassificationAudio TaggingRetrievalTransfer Learning+1