paper-with-me

홈 › Papers

NIESR: Nuisance Invariant End-to-end Speech Recognition

2019-07-07 · I-Hung Hsu, Ayush Jaiswal, Premkumar Natarajan

Deep neural network models for speech recognition have achieved great success recently, but they can learn incorrect associations between the target and nuisance factors of speech (e.g., speaker identities, background noise, etc.), which can lead to overfitting. While several methods have been proposed to tackle this problem, existing methods incorporate additional information about nuisance factors during training to develop invariant models. However, enumeration of all possible nuisance factors in speech data and the collection of their annotations is difficult and expensive. We present a robust training scheme for end-to-end speech recognition that adopts an unsupervised adversarial invariance induction framework to separate out essential factors for speech-recognition from nuisances without using any supplementary labels besides the transcriptions. Experiments show that the speech recognition model trained with the proposed training scheme achieves relative improvements of 5.48% on WSJ0, 6.16% on CHiME3, and 6.61% on TIMIT dataset over the base model. Additionally, the proposed method achieves a relative improvement of 14.44% on the combined WSJ0+CHiME3 dataset.

📄 PDF Abstract BibTeX arXiv:1907.03233

Code (1)

ihungalexhsu/NIESR-code 공식 구현 pytorch

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Rate-Invariant Analysis of Trajectories on Riemannian Manifolds with Application in Visual Speech Recognition

2014-06-01 · CVPR 2014 6 · Jingyong Su, Anuj Srivastava, Fillipe D. M. de Souza, Sudeep Sarkar

In statistical analysis of video sequences for speech recognition, and more generally activity recognition, it is natural to treat temporal evolutions of features as trajectories on Riemannian manifolds. However, differe…

Activity RecognitionClassificationGeneral Classificationspeech-recognition+2

Nuisance-Label Supervision: Robustness Improvement by Free Labels

2021-10-14 · Xinyue Wei, Weichao Qiu, Yi Zhang, Zihao Xiao 외

In this paper, we present a Nuisance-label Supervision (NLS) module, which can make models more robust to nuisance factor variations. Nuisance factors are those irrelevant to a task, and an ideal model should be invarian…

Action RecognitionActivity RecognitionData Augmentation

Untangling in Invariant Speech Recognition

2020-03-03 · NeurIPS 2019 12 · Cory Stephenson, Jenelle Feather, Suchismita Padhy, Oguz Elibol 외

Encouraged by the success of deep neural networks on a variety of visual tasks, much theoretical and experimental work has been aimed at understanding and interpreting how vision networks operate. Meanwhile, deep neural …

speech-recognitionSpeech Recognition

A Supervised Low-Rank Method for Learning Invariant Subspaces

2015-12-01 · ICCV 2015 12 · Farzad Siyahjani, Ranya Almohsen, Sinan Sabri, Gianfranco Doretto

Sparse representation and low-rank matrix decomposition approaches have been successfully applied to several computer vision problems. They build a generative representation of the data, which often requires complex trai…

Face RecognitionMetric LearningRobust classification

Unsupervised Adaptation with Interpretable Disentangled Representations for Distant Conversational Speech Recognition

2018-06-13 · Wei-Ning Hsu, Hao Tang, James Glass

The current trend in automatic speech recognition is to leverage large amounts of labeled data to train supervised neural network models. Unfortunately, obtaining data for a wide range of domains to train robust models c…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition