Benchmarking Representations for Speech, Music, and Acoustic Events
Limited diversity in standardized benchmarks for evaluating audio representation learning (ARL) methods may hinder systematic comparison of current methods' capabilities. We present ARCH, a comprehensive benchmark for evaluating ARL methods on diverse audio classification domains, covering acoustic events, music, and speech. ARCH comprises 12 datasets, that allow us to thoroughly assess pre-trained SSL models of different sizes. ARCH streamlines benchmarking of ARL techniques through its unified access to a wide range of domains and its ability to readily incorporate new datasets and models. To address the current lack of open-source, pre-trained models for non-speech audio, we also release new pre-trained models that demonstrate strong performance on non-speech datasets. We argue that the presented wide-ranging evaluation provides valuable insights into state-of-the-art ARL methods, and is useful to pinpoint promising research directions.
Code (1)
Tasks
Audio ClassificationBenchmarkingDiversityRepresentation LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and…
Domain AdaptationDomain GeneralizationEmotion RecognitionMusic Emotion Recognition+3The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks
The cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research. Recent efforts have mainly focused on separating speech from no…
Audio Source SeparationCross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music
Modern audio systems universally employ mel-scale representations derived from 1940s Western psychoacoustic studies, potentially encoding cultural biases that create systematic performance disparities. We present a compr…
Acoustic Scene ClassificationSpeech RecognitionNatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
Large language models (LLMs) prompted with text and audio represent the state of the art in various auditory tasks, including speech, music, and general audio, showing emergent abilities on unseen tasks. However, these c…
zero-shot-classificationZero-Shot LearningJoint Blind Room Acoustic Characterization From Speech And Music Signals Using Convolutional Recurrent Neural Networks
Acoustic environment characterization opens doors for sound reproduction innovations, smart EQing, speech enhancement, hearing aids, and forensics. Reverberation time, clarity, and direct-to-reverberant ratio are acousti…
parameter estimationRoom Impulse Response (RIR)Speech Enhancement