paper-with-me

Papers

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

2026-03-02 · Zijian Yang, Jörg Barkoczi, Ralf Schlüter, Hermann Ney arxiv

Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how classification error relates to candidate training objectives, we develop a theoretical framework for unsupervised speech recognition grounded in classification error bounds. We introduce two conditions under which unsupervised speech recognition is possible. The necessity of these conditions are also discussed. Under these conditions, we derive a classification error bound for unsupervised speech recognition and validate this bound in simulations. Motivated by this bound, we propose a single-stage sequence-level cross-entropy loss for unsupervised speech recognition.

📄 PDF Abstract BibTeX arXiv:2603.02285

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Progressive Joint Modeling in Unsupervised Single-channel Overlapped Speech Recognition

2017-07-21 · Zhehuai Chen, Jasha Droppo, Jinyu Li, Wayne Xiong

Unsupervised single-channel overlapped speech recognition is one of the hardest problems in automatic speech recognition (ASR). Permutation invariant training (PIT) is a state of the art model-based approach, which appli…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Towards Unsupervised Automatic Speech Recognition Trained by Unaligned Speech and Text only

2018-03-29 · Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang, Hung-Yi Lee

Automatic speech recognition (ASR) has been widely researched with supervised approaches, while many low-resourced languages lack audio-text aligned data, and supervised methods cannot be applied on them. In this work,…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Almost Unsupervised Text to Speech and Automatic Speech Recognition

2019-05-13 · Yi Ren, Xu Tan, Tao Qin, Sheng Zhao 외

Text to speech (TTS) and automatic speech recognition (ASR) are two dual tasks in speech processing and both achieve impressive performance thanks to the recent advance in deep learning and large amount of aligned speech…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingLanguage Modeling+5

Unsupervised pre-training for sequence to sequence speech recognition

2019-10-28 · Zhiyun Fan, Shiyu Zhou, Bo Xu

This paper proposes a novel approach to pre-train encoder-decoder sequence-to-sequence (seq2seq) model with unpaired speech and transcripts respectively. Our pre-training method is divided into two stages, named acoustic…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderSequence-To-Sequence Speech Recognition+5

Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization

2024-01-13 · A F M Saif, Xiaodong Cui, Han Shen, Songtao Lu 외

In this paper, we present a novel bilevel optimization-based training approach to training acoustic models for automatic speech recognition (ASR) tasks that we term {bi-level joint unsupervised and supervised training (B…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Bilevel Optimizationspeech-recognition+1