paper-with-me

홈 › Papers

Content-Context Factorized Representations for Automated Speech Recognition

2022-05-19 · David M. Chan, Shalini Ghosh

Deep neural networks have largely demonstrated their ability to perform automated speech recognition (ASR) by extracting meaningful features from input audio frames. Such features, however, may consist not only of information about the spoken language content, but also may contain information about unnecessary contexts such as background noise and sounds or speaker identity, accent, or protected attributes. Such information can directly harm generalization performance, by introducing spurious correlations between the spoken words and the context in which such words were spoken. In this work, we introduce an unsupervised, encoder-agnostic method for factoring speech-encoder representations into explicit content-encoding representations and spurious context-encoding representations. By doing so, we demonstrate improved performance on standard ASR benchmarks, as well as improved performance in both real-world and artificially noisy ASR scenarios.

📄 PDF Abstract BibTeX arXiv:2205.09872

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Speech Disorder Classification Using Extended Factorized Hierarchical Variational Auto-encoders

2021-06-14 · Jinzi Qi, Hugo Van hamme

Objective speech disorder classification for speakers with communication difficulty is desirable for diagnosis and administering therapy. With the current state of speech technology, it is evident to propose neural netwo…

ClassificationDisentanglementRepresentation LearningSentence

Streaming Neural Speech Codecs through Time-Invariant Representations

2026-07-06 · Kélian Estève, Salima Mhdaffar, Mickael Rouvier, Richard Dufour 외 arxiv

Neural speech codecs are increasingly used as intermediate representations in codec-based speech generation systems. TiCodec introduces a factorized representation that separates time-varying speech content from time-inv…

Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data

2017-09-22 · NeurIPS 2017 12 · Wei-Ning Hsu, Yu Zhang, James Glass

We present a factorized hierarchical variational autoencoder, which learns disentangled and interpretable representations from sequential data without supervision. Specifically, we exploit the multi-scale nature of infor…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Verificationspeech-recognition+1

NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

2024-03-05 · Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan 외

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.…

QuantizationSpeech Synthesistext-to-speechText to Speech

Learning Subject-Invariant Representations from Speech-Evoked EEG Using Variational Autoencoders

2022-07-01 · Lies Bollens, Tom Francart, Hugo Van hamme

The electroencephalogram (EEG) is a powerful method to understand how the brain processes speech. Linear models have recently been replaced for this purpose with deep neural networks and yield promising results. In relat…

ClassificationEEGElectroencephalogram (EEG)