paper-with-me

Papers

Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing

2025-01-01 · Gaofeng Cheng, Haitian Lu, Chengxu Yang, Xuyang Wang, Ta Li, Yonghong Yan

Effectively distinguishing the pronunciation correlations between different written texts is a significant issue in linguistic acoustics. Traditionally, such pronunciation correlations are obtained through manually designed pronunciation lexicons. In this paper, we propose a data-driven method to automatically acquire these pronunciation correlations, called automatic text pronunciation correlation (ATPC). The supervision required for this method is consistent with the supervision needed for training end-to-end automatic speech recognition (E2E-ASR) systems, i.e., speech and corresponding text annotations. First, the iteratively-trained timestamp estimator (ITSE) algorithm is employed to align the speech with their corresponding annotated text symbols. Then, a speech encoder is used to convert the speech into speech embeddings. Finally, we compare the speech embeddings distances of different text symbols to obtain ATPC. Experimental results on Mandarin show that ATPC enhances E2E-ASR performance in contextual biasing and holds promise for dialects or languages lacking artificial pronunciation lexicons.

📄 PDF Abstract BibTeX arXiv:2501.00804

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Automatic Pronunciation Assessment using Self-Supervised Speech Representation Learning

2022-04-08 · Eesung Kim, Jae-Jin Jeon, Hyeji Seo, Hoon Kim

Self-supervised learning (SSL) approaches such as wav2vec 2.0 and HuBERT models have shown promising results in various downstream tasks in the speech community. In particular, speech representations learned by SSL model…

Representation LearningSelf-Supervised LearningSpeech Representation Learning

Fine-Tuning Self-Supervised Learning Models for End-to-End Pronunciation Scoring

2023-09-19 · IEEE Access 2023 9 · Ahmed I. Zahran, Aly A. Fahmy, Khaled T. Wassif, Hanaa Bayomi

Automatic pronunciation assessment models are regularly used in language learning applications. Common methodologies for pronunciation assessment use feature-based approaches, such as the Goodness-of-Pronunciation (GOP) …

Feature EngineeringPhone-level pronunciation scoringPhoneme RecognitionSelf-Supervised Learning+3

A Generative Model of a Pronunciation Lexicon for Hindi

2017-05-06 · Pramod Pandey, Somnath Roy

Voice browser applications in Text-to- Speech (TTS) and Automatic Speech Recognition (ASR) systems crucially depend on a pronunciation lexicon. The present paper describes the model of pronunciation lexicon of Hindi deve…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Multi-granularity Interactive Attention Framework for Residual Hierarchical Pronunciation Assessment

2026-01-05 · Hong Han, Hao-Chen Pei, Zhao-Zheng Nie, Xin Luo 외 arxiv

Automatic pronunciation assessment plays a crucial role in computer-assisted pronunciation training systems. Due to the ability to perform multiple pronunciation tasks simultaneously, multi-aspect multi-granularity pronu…

Zero-Shot Automatic Pronunciation Assessment

2023-05-31 · Hongfu Liu, Mingqian Shi, Ye Wang

Automatic Pronunciation Assessment (APA) is vital for computer-assisted language learning. Prior methods rely on annotated speech-text data to train Automatic Speech Recognition (ASR) models or speech-score data to train…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringregression+2