paper-with-me

Papers

Unsupervised Speech Recognition via Segmental Empirical Output Distribution Matching

2018-12-23 · ICLR 2019 5 · Chih-Kuan Yeh, Jianshu Chen, Chengzhu Yu, Dong Yu

We consider the problem of training speech recognition systems without using any labeled data, under the assumption that the learner can only access to the input utterances and a phoneme language model estimated from a non-overlapping corpus. We propose a fully unsupervised learning algorithm that alternates between solving two sub-problems: (i) learn a phoneme classifier for a given set of phoneme segmentation boundaries, and (ii) refining the phoneme boundaries based on a given classifier. To solve the first sub-problem, we introduce a novel unsupervised cost function named Segmental Empirical Output Distribution Matching, which generalizes the work in (Liu et al., 2017) to segmental structures. For the second sub-problem, we develop an approximate MAP approach to refining the boundaries obtained from Wang et al. (2017). Experimental results on TIMIT dataset demonstrate the success of this fully unsupervised phoneme recognition system, which achieves a phone error rate (PER) of 41.6%. Although it is still far away from the state-of-the-art supervised systems, we show that with oracle boundaries and matching language model, the PER could be improved to 32.5%.This performance approaches the supervised system of the same model architecture, demonstrating the great potential of the proposed method.

📄 PDF Abstract BibTeX arXiv:1812.09323

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingPhoneme Recognitionspeech-recognitionSpeech RecognitionUnsupervised Speech Recognition

Similar Papers 제목 키워드 기반

A segmental framework for fully-unsupervised large-vocabulary speech recognition

2016-06-22 · Herman Kamper, Aren Jansen, Sharon Goldwater

Zero-resource speech technology is a growing research area that aims to develop methods for speech processing in the absence of transcriptions, lexicons, or language modelling text. Early term discovery systems focused o…

Language ModellingSpeech RecognitionUnsupervised Speech Recognition

REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR

2024-02-06 · Liang-Hsuan Tseng, En-Pei Hu, Cheng-Han Chiang, Yuan Tseng 외

Unsupervised automatic speech recognition (ASR) aims to learn the mapping between the speech signal and its corresponding textual transcription without the supervision of paired speech-text data. A word/phoneme in the sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Segmentationspeech-recognition+1

Efficient Segmental Cascades for Speech Recognition

2016-08-02 · Hao Tang, Weiran Wang, Kevin Gimpel, Karen Livescu

Discriminative segmental models offer a way to incorporate flexible feature functions into speech recognition. However, their appeal has been limited by their computational requirements, due to the large number of possib…

speech-recognitionSpeech Recognition

Multitask Learning with CTC and Segmental CRF for Speech Recognition

2017-02-21 · Liang Lu, Lingpeng Kong, Chris Dyer, Noah A. Smith

Segmental conditional random fields (SCRFs) and connectionist temporal classification (CTC) are two sequence labeling methods used for end-to-end training of speech recognition models. Both models define a transcription …

speech-recognitionSpeech Recognition

Automatic recognition of suprasegmentals in speech

2021-08-02 · Jiahong Yuan, Neville Ryant, Xingyu Cai, Kenneth Church 외

This study reports our efforts to improve automatic recognition of suprasegmentals by fine-tuning wav2vec 2.0 with CTC, a method that has been successful in automatic speech recognition. We demonstrate that the method ca…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phoneme Recognitionspeech-recognition+1