paper-with-me

Papers

Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance

2024-08-12 · Manuel Milling, Shuo Liu, Andreas Triantafyllopoulos, Ilhan Aslan, Björn W. Schuller

Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module, which can be developed independently, is explicitly used at the front-end of the target audio applications. In this paper, we present an end-to-end learning solution to jointly optimise the models for audio enhancement (AE) and the subsequent applications. To guide the optimisation of the AE module towards a target application, and especially to overcome difficult samples, we make use of the sample-wise performance measure as an indication of sample importance. In experiments, we consider four representative applications to evaluate our training paradigm, i.e., ASR, speech command recognition (SCR), speech emotion recognition (SER), and ASC. These applications are associated with speech and non-speech tasks concerning semantic and non-semantic features, transient and global information, and the experimental results indicate that our proposed approach can considerably boost the noise robustness of the models, especially at low signal-to-noise ratios (SNRs), for a wide range of computer audition tasks in everyday-life noisy environments.

📄 PDF Abstract BibTeX arXiv:2408.06264

Code (0)

등록된 구현이 없습니다.

Tasks

Acoustic Scene ClassificationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionScene ClassificationSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

AE An autoencoder is a type of artificial neural network used to learn efficient data codings in an unsupervised manner. The aim of an autoencoder is to learn a representation…

Similar Papers 제목 키워드 기반

Climate Change & Computer Audition: A Call to Action and Overview on Audio Intelligence to Help Save the Planet

2022-03-10 · Björn W. Schuller, Alican Akman, Yi Chang, Harry Coppock 외

Among the seventeen Sustainable Development Goals (SDGs) proposed within the 2030 Agenda and adopted by all the United Nations member states, the 13$^{th}$ SDG is a call for action to combat climate change for a better w…

autrainer: A Modular and Extensible Deep Learning Toolkit for Computer Audition Tasks

2024-12-16 · Simon Rampp, Andreas Triantafyllopoulos, Manuel Milling, Björn W. Schuller

This work introduces the key operating principles for autrainer, our new deep learning training framework for computer audition tasks. autrainer is a PyTorch-based toolkit that allows for rapid, reproducible, and easily …

OtoWorld: Towards Learning to Separate by Learning to Move

2020-07-12 · Omkar Ranadive, Grant Gasser, David Terpay, Prem Seetharaman

We present OtoWorld, an interactive environment in which agents must learn to listen in order to solve navigational tasks. The purpose of OtoWorld is to facilitate reinforcement learning research in computer audition, wh…

Audio Source SeparationNavigateOpenAI Gym

DroneAudioset: An Audio Dataset for Drone-based Search and Rescue

2025-10-17 · Chitralekha Gupta, Soundarya Ramesh, Praveen Sasikumar, Kian Peen Yeo 외 arxiv

Unmanned Aerial Vehicles (UAVs) or drones, are increasingly used in search and rescue missions to detect human presence. Existing systems primarily leverage vision-based methods which are prone to fail under low-visibili…

Audio Self-supervised Learning: A Survey

2022-03-02 · Shuo Liu, Adria Mallol-Ragolta, Emilia Parada-Cabeleiro, Kun Qian 외

Inspired by the humans' cognitive ability to generalise knowledge and skills, Self-Supervised Learning (SSL) targets at discovering general representations from large-scale data without requiring human annotations, which…

Self-Supervised LearningSurvey