paper-with-me

Papers

A consolidated view of loss functions for supervised deep learning-based speech enhancement

2020-09-25

Deep learning-based speech enhancement for real-time applications recently made large advancements. Due to the lack of a tractable perceptual optimization target, many myths around training losses emerged, whereas the contribution to success of the loss functions in many cases has not been investigated isolated from other factors such as network architecture, features, or training procedures. In this work, we investigate a wide variety of loss spectral functions for a recurrent neural network architecture suitable to operate in online frame-by-frame processing. We relate magnitude-only with phase-aware losses, ratios, correlation metrics, and compressed metrics. Our results reveal that combining magnitude-only with phase-aware objectives always leads to improvements, even when the phase is not enhanced. Furthermore, using compressed spectral values also yields a significant improvement. On the other hand, phase-sensitive improvement is best achieved by linear domain losses such as mean absolute error.

📄 PDF Abstract BibTeX arXiv:2009.12286

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Perceive and predict: self-supervised speech representation based loss functions for speech enhancement

2023-01-11 · George Close, William Ravenscroft, Thomas Hain, Stefan Goetze

Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the d…

Speech Enhancement

The Effect of Spoken Language on Speech Enhancement using Self-Supervised Speech Representation Loss Functions

2023-07-27 · George Close, Thomas Hain, Stefan Goetze

Recent work in the field of speech enhancement (SE) has involved the use of self-supervised speech representations (SSSRs) as feature transformations in loss functions. However, in prior work, very little attention has b…

Speech Enhancement

Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition

2024-12-11 · Xiaodong Cui, A F M Saif, Songtao Lu, Lisha Chen 외

In this paper, we propose a bilevel joint unsupervised and supervised training (BL-JUST) framework for automatic speech recognition. Compared to the conventional pre-training and fine-tuning strategy which is a disconnec…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

In Data We Trust: A Critical Analysis of Hate Speech Detection Datasets

2020-11-01 · EMNLP (ALW) 2020 11 · Kosisochukwu Madukwe, Xiaoying Gao, Bing Xue

Recently, a few studies have discussed the limitations of datasets collected for the task of detecting hate speech from different viewpoints. We intend to contribute to the conversation by providing a consolidated overvi…

Hate Speech Detection

Single-channel speech enhancement using learnable loss mixup

2023-12-20 · Oscar Chang, Dung N. Tran, Kazuhito Koishida

Generalization remains a major problem in supervised learning of single-channel speech enhancement. In this work, we propose learnable loss mixup (LLM), a simple and effortless training diagram, to improve the generaliza…

Speech Enhancement