paper-with-me

홈 › Papers

Benchmarking Representations for Speech, Music, and Acoustic Events

2024-05-02 · Moreno La Quatra, Alkis Koudounas, Lorenzo Vaiani, Elena Baralis, Luca Cagliero, Paolo Garza, Sabato Marco Siniscalchi

Limited diversity in standardized benchmarks for evaluating audio representation learning (ARL) methods may hinder systematic comparison of current methods' capabilities. We present ARCH, a comprehensive benchmark for evaluating ARL methods on diverse audio classification domains, covering acoustic events, music, and speech. ARCH comprises 12 datasets, that allow us to thoroughly assess pre-trained SSL models of different sizes. ARCH streamlines benchmarking of ARL techniques through its unified access to a wide range of domains and its ability to readily incorporate new datasets and models. To address the current lack of open-source, pre-trained models for non-speech audio, we also release new pre-trained models that demonstrate strong performance on non-speech datasets. We argue that the presented wide-ranging evaluation provides valuable insights into state-of-the-art ARL methods, and is useful to pinpoint promising research directions.

📄 PDF Abstract BibTeX arXiv:2405.00934

Code (1)

MorenoLaQuatra/ARCH 공식 구현 pytorch

Tasks

Audio ClassificationBenchmarkingDiversityRepresentation Learning

Methods 이 논문이 사용한 방법론

ARCH Animatable Reconstruction of Clothed Humans is an end-to-end framework for accurate reconstruction of animation-ready 3D clothed humans from a monocular image. ARCH is a…

Similar Papers 제목 키워드 기반

Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations

2024-09-26 · Yujia Sun, Zeyu Zhao, Korin Richmond, Yuanchao Li

Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and…

Domain AdaptationDomain GeneralizationEmotion RecognitionMusic Emotion Recognition+3

The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks

2021-10-19 · Darius Petermann, Gordon Wichern, Zhong-Qiu Wang, Jonathan Le Roux

The cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research. Recent efforts have mainly focused on separating speech from no…

Audio Source Separation

Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music

2026-04-12 · Shivam Chauhan, Ajay Pundhir arxiv

Modern audio systems universally employ mel-scale representations derived from 1940s Western psychoacoustic studies, potentially encoding cultural biases that create systematic performance disparities. We present a compr…

Acoustic Scene ClassificationSpeech Recognition

NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics

2024-11-11 · David Robinson, Marius Miron, Masato Hagiwara, Olivier Pietquin

Large language models (LLMs) prompted with text and audio represent the state of the art in various auditory tasks, including speech, music, and general audio, showing emergent abilities on unseen tasks. However, these c…

zero-shot-classificationZero-Shot Learning

Joint Blind Room Acoustic Characterization From Speech And Music Signals Using Convolutional Recurrent Neural Networks

2020-10-21 · Paul Callens, Milos Cernak

Acoustic environment characterization opens doors for sound reproduction innovations, smart EQing, speech enhancement, hearing aids, and forensics. Reverberation time, clarity, and direct-to-reverberant ratio are acousti…

parameter estimationRoom Impulse Response (RIR)Speech Enhancement