paper-with-me

Papers

PAAPLoss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement

2023-02-16 · Muqiao Yang, Joseph Konan, David Bick, Yunyang Zeng, Shuo Han, Anurag Kumar, Shinji Watanabe, Bhiksha Raj

Despite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in perceptual quality, by using domain knowledge of acoustic-phonetics. We identify temporal acoustic parameters -- such as spectral tilt, spectral flux, shimmer, etc. -- that are non-differentiable, and we develop a neural network estimator that can accurately predict their time-series values across an utterance. We also model phoneme-specific weights for each feature, as the acoustic parameters are known to show different behavior in different phonemes. We can add this criterion as an auxiliary loss to any model that produces speech, to optimize speech outputs to match the values of clean speech in these features. Experimentally we show that it improves speech enhancement workflows in both time-domain and time-frequency domain, as measured by standard evaluation metrics. We also provide an analysis of phoneme-dependent improvement on acoustic parameters, demonstrating the additional interpretability that our method provides. This analysis can suggest which features are currently the bottleneck for improvement.

📄 PDF Abstract BibTeX arXiv:2302.08095

Code (2)

muqiaoy/PAAP 공식 구현 pytorch
yunyangzeng/taploss pytorch

Tasks

Speech EnhancementTime SeriesTime Series Analysis

Similar Papers 제목 키워드 기반

Learning acoustic word embeddings with phonetically associated triplet network

2018-11-07 · Hyungjun Lim, Younggwan Kim, Youngmoon Jung, Myunghun Jung 외

Previous researches on acoustic word embeddings used in query-by-example spoken term detection have shown remarkable performance improvements when using a triplet network. However, the triplet network is trained using on…

TripletWord Embeddings

Common Phone: A Multilingual Dataset for Robust Acoustic Modelling

2022-01-15 · LREC 2022 6 · Philipp Klumpp, Tomás Arias-Vergara, Paula Andrea Pérez-Toro, Elmar Nöth 외

Current state of the art acoustic models can easily comprise more than 100 million parameters. This growing complexity demands larger training datasets to maintain a decent generalization of the final decision function. …

Acoustic Modellingparameter estimation

Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription

2018-02-01 · Yu Wang, Xie Chen, Mark Gales, Anton Ragni 외

State-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Filter-based Discriminative Autoencoders for Children Speech Recognition

2022-04-01 · Chiang-Lin Tai, Hung-Shin Lee, Yu Tsao, Hsin-Min Wang

Children speech recognition is indispensable but challenging due to the diversity of children's speech. In this paper, we propose a filter-based discriminative autoencoder for acoustic modeling. To filter out the influen…

DecoderDiversityDomain Adaptationspeech-recognition+1

A Corpus for Large-Scale Phonetic Typology

2020-05-28 · ACL 2020 6 · Elizabeth Salesky, Eleanor Chodroff, Tiago Pimentel, Matthew Wiesner 외

A major hurdle in data-driven research on typology is having sufficient data in many languages to draw meaningful conclusions. We present VoxClamantis v1.0, the first large-scale corpus for phonetic typology, with aligne…