paper-with-me

홈 › Papers

Predicting within and across language phoneme recognition performance of self-supervised learning speech pre-trained models

2022-06-24 · Hang Ji, Tanvina Patel, Odette Scharenborg

In this work, we analyzed and compared speech representations extracted from different frozen self-supervised learning (SSL) speech pre-trained models on their ability to capture articulatory features (AF) information and their subsequent prediction of phone recognition performance for within and across language scenarios. Specifically, we compared CPC, wav2vec 2.0, and HuBert. First, frame-level AF probing tasks were implemented. Subsequently, phone-level end-to-end ASR systems for phoneme recognition tasks were implemented, and the performance on the frame-level AF probing task and the phone accuracy were correlated. Compared to the conventional speech representation MFCC, all SSL pre-trained speech representations captured more AF information, and achieved better phoneme recognition performance within and across languages, with HuBert performing best. The frame-level AF probing task is a good predictor of phoneme recognition performance, showing the importance of capturing AF information in the speech representations. Compared with MFCC, in the within-language scenario, the performance of these SSL speech pre-trained models on AF probing tasks achieved a maximum relative increase of 34.4%, and it resulted in the lowest PER of 10.2%. In the cross-language scenario, the maximum relative increase of 26.7% also resulted in the lowest PER of 23.0%.

📄 PDF Abstract BibTeX arXiv:2206.12489

Code (1)

karenmars/is22code 공식 구현

Tasks

Phoneme RecognitionSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Zero-shot Learning for Speech Recognition with Universal Phonetic Model

2018-09-27 · Xinjian Li, Siddharth Dalmia, David R. Mortensen, Florian Metze 외

There are more than 7,000 languages in the world, but due to the lack of training sets, only a small number of them have speech recognition systems. Multilingual speech recognition provides a solution if at least some au…

speech-recognitionSpeech RecognitionZero-Shot Learning

Multilingual Speech Recognition for Low-Resource Indian Languages using Multi-Task conformer

2021-08-22 · Krishna D N

Transformers have recently become very popular for sequence-to-sequence applications such as machine translation and speech recognition. In this work, we propose a multi-task learning-based transformer model for low-reso…

DecoderMachine TranslationMulti-Task LearningPhoneme Recognition+3

Optimizing Two-Pass Cross-Lingual Transfer Learning: Phoneme Recognition and Phoneme to Grapheme Translation

2023-12-06 · Wonjun Lee, Gary Geunbae Lee, Yunsu Kim

This research optimizes two-pass cross-lingual transfer learning in low-resource languages by enhancing phoneme recognition and phoneme-to-grapheme translation models. Our approach optimizes these two stages to improve s…

Cross-Lingual TransferPhoneme Recognitionspeech-recognitionSpeech Recognition+1

Segment Boundary Detection via Class Entropy Measurements in Connectionist Phoneme Recognition

2024-01-11 · Giampiero Salvi

This article investigates the possibility to use the class entropy of the output of a connectionist phoneme recogniser to predict time boundaries between phonetic classes. The rationale is that the value of the entropy s…

Boundary DetectionPhoneme Recognition

Metric Learning for Phoneme Perception

2018-09-20 · Yair Lakretz, Gal Chechik, Evan-Gary Cohen, Alessandro Treves 외

Metric functions for phoneme perception capture the similarity structure among phonemes in a given language and therefore play a central role in phonology and psycho-linguistics. Various phenomena depend on phoneme simil…

Metric Learning