paper-with-me

Papers

The Effect of Spoken Language on Speech Enhancement using Self-Supervised Speech Representation Loss Functions

2023-07-27 · George Close, Thomas Hain, Stefan Goetze

Recent work in the field of speech enhancement (SE) has involved the use of self-supervised speech representations (SSSRs) as feature transformations in loss functions. However, in prior work, very little attention has been paid to the relationship between the language of the audio used to train the self-supervised representation and that used to train the SE system. Enhancement models trained using a loss function which incorporates a self-supervised representation that shares exactly the language of the noisy data used to train the SE system show better performance than those which do not match exactly. This may lead to enhancement systems which are language specific and as such do not generalise well to unseen languages, unlike models trained using traditional spectrogram or time domain loss functions. In this work, SE models are trained and tested on a number of different languages, with self-supervised representations which themselves are trained using different language combinations and with differing network structures as loss function representations. These models are then tested across unseen languages and their performances are analysed. It is found that the training language of the self-supervised representation appears to have a minor effect on enhancement performance, the amount of training data of a particular language, however, greatly affects performance.

📄 PDF Abstract BibTeX arXiv:2307.14502

Code (1)

leto19/commonvoice-demand 공식 구현

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Spoken Speech Enhancement using EEG

2019-09-13 · Gautam Krishna, Co Tran, Yan Han, Mason Carnahan 외

In this paper we demonstrate spoken speech enhancement using electroencephalography (EEG) signals using a generative adversarial network (GAN) based model, gated recurrent unit (GRU) regression based model, temporal conv…

EEGElectroencephalogram (EEG)Generative Adversarial Networkregression+1

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

2021-10-14 · ACL 2022 5 · Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang 외

Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for self-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderQuantization+7

Enhancements in statistical spoken language translation by de-normalization of ASR results

2015-11-18 · Agnieszka Wołk, Krzysztof Wołk, Krzysztof Marasek

Spoken language translation (SLT) has become very important in an increasingly globalized world. Machine translation (MT) for automatic speech recognition (ASR) systems is a major challenge of great interest. This resear…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSegmentation+5

DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding

2024-06-13 · Suwon Shon, Kwangyoun Kim, Yi-Te Hsu, Prashant Sridhar 외

The integration of pre-trained text-based large language models (LLM) with speech input has enabled instruction-following capabilities for diverse speech tasks. This integration requires the use of a speech encoder, a sp…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+2

A Systematic Comparison of Phonetic Aware Techniques for Speech Enhancement

2022-06-22 · Or Tal, Moshe Mandel, Felix Kreuk, Yossi Adi

Speech enhancement has seen great improvement in recent years using end-to-end neural networks. However, most models are agnostic to the spoken phonetic content. Recently, several studies suggested phonetic-aware speech …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Model OptimizationSelf-Supervised Learning+2