paper-with-me

Papers

Improved far-field speech recognition using Joint Variational Autoencoder

2022-04-24 · Shashi Kumar, Shakti P. Rath, Abhishek Pandey

Automatic Speech Recognition (ASR) systems suffer considerably when source speech is corrupted with noise or room impulse responses (RIR). Typically, speech enhancement is applied in both mismatched and matched scenario training and testing. In matched setting, acoustic model (AM) is trained on dereverberated far-field features while in mismatched setting, AM is fixed. In recent past, mapping speech features from far-field to close-talk using denoising autoencoder (DA) has been explored. In this paper, we focus on matched scenario training and show that the proposed joint VAE based mapping achieves a significant improvement over DA. Specifically, we observe an absolute improvement of 2.5% in word error rate (WER) compared to DA based enhancement and 3.96% compared to AM trained directly on far-field filterbank features.

📄 PDF Abstract BibTeX arXiv:2204.11286

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingSpeech Enhancementspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

AM 설명 없음
Denoising Autoencoder A Denoising Autoencoder is a modification on the autoencoder to prevent the network learning the identity function.…

Similar Papers 제목 키워드 기반

Constrained Variational Autoencoder for improving EEG based Speech Recognition Systems

2020-06-01 · Gautam Krishna, Co Tran, Mason Carnahan, Ahmed Tewfik

In this paper we introduce a recurrent neural network (RNN) based variational autoencoder (VAE) model with a new constrained loss function that can generate more meaningful electroencephalography (EEG) features from raw …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Brain Computer InterfaceEEG+3

Learning Speech Representations with Variational Predictive Coding

2025-12-31 · Sung-Lin Yeh, Peter Bell, Hao Tang arxiv

Despite being the best known objective for learning speech representations, the HuBERT objective has not been further developed and improved. We argue that it is the lack of an underlying principle that stalls the develo…

Speaker RecognitionSpeech Recognition

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning

2019-08-13 · Pavel Denisov, Ngoc Thang Vu

This paper presents our latest investigation on end-to-end automatic speech recognition (ASR) for overlapped speech. We propose to train an end-to-end system conditioned on speaker embeddings and further improved by tran…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding

2025-01-13 · Jiliang Hu, Zuchao Li, Mengjia Shen, Haojun Ai 외

Spoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved great success. However, This method is not…

Automatic Speech Recognitionintent-classificationIntent ClassificationNER+3

Multichannel End-to-end Speech Recognition

2017-03-14 · ICML 2017 8 · Tsubasa Ochiai, Shinji Watanabe, Takaaki Hori, John R. Hershey

The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent enco…

DecoderLanguage ModelingLanguage ModellingSpeech Enhancement+2