paper-with-me

Papers

Generative error correction for code-switching speech recognition using large language models

2023-10-17 · Chen Chen, Yuchen Hu, Chao-Han Huck Yang, Hexin Liu, Sabato Marco Siniscalchi, Eng Siong Chng

Code-switching (CS) speech refers to the phenomenon of mixing two or more languages within the same sentence. Despite the recent advances in automatic speech recognition (ASR), CS-ASR is still a challenging task ought to the grammatical structure complexity of the phenomenon and the data scarcity of specific training corpus. In this work, we propose to leverage large language models (LLMs) and lists of hypotheses generated by an ASR to address the CS problem. Specifically, we first employ multiple well-trained ASR models for N-best hypotheses generation, with the aim of increasing the diverse and informative elements in the set of hypotheses. Next, we utilize the LLMs to learn the hypotheses-to-transcription (H2T) mapping by adding a trainable low-rank adapter. Such a generative error correction (GER) method directly predicts the accurate transcription according to its expert linguistic knowledge and N-best hypotheses, resulting in a paradigm shift from the traditional language model rescoring or error correction techniques. Experimental evidence demonstrates that GER significantly enhances CS-ASR accuracy, in terms of reduced mixed error rate (MER). Furthermore, LLMs show remarkable data efficiency for H2T learning, providing a potential solution to the data scarcity problem of CS-ASR in low-resource languages.

📄 PDF Abstract BibTeX arXiv:2310.13013

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModellingSentencespeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
GCN A Graph Convolutional Network, or GCN, is an approach for semi-supervised learning on graph-structured data. It is based on an efficient variant of [convolutional neural…
GER We present a novel classifier network called STEP, to classify perceived human emotion from gaits, based on a Spatial Temporal Graph Convolutional Network…

Similar Papers 제목 키워드 기반

Aligning Speech to Languages to Enhance Code-switching Speech Recognition

2024-03-09 · Hexin Liu, Xiangyu Zhang, Haoyang Zhang, Leibny Paola Garcia 외

Code-switching (CS) refers to the switching of languages within a speech signal and results in language confusion for automatic speech recognition (ASR). To address language confusion, we introduce a novel language align…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderLanguage Identification+4

Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition

2023-10-10 · Srijith Radhakrishnan, Chao-Han Huck Yang, Sumeer Ahmad Khan, Rohit Kumar 외

We introduce a new cross-modal fusion technique designed for generative error correction in automatic speech recognition (ASR). Our methodology leverages both acoustic information and external linguistic representations …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Can Generative Large Language Models Perform ASR Error Correction?

2023-07-09 · Rao Ma, Mengjie Qian, Potsawee Manakul, Mark Gales 외

ASR error correction is an interesting option for post processing speech recognition system outputs. These error correction models are usually trained in a supervised fashion using the decoding results of a target ASR sy…

Decoderspeech-recognitionSpeech Recognition

ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation

2021-12-12 · LREC 2022 6 · Holy Lovenia, Samuel Cahyawijaya, Genta Indra Winata, Peng Xu 외

Code-switching is a speech phenomenon occurring when a speaker switches language during a conversation. Despite the spontaneous nature of code-switching in conversational spoken language, most existing works collect code…

CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR

2025-05-24 · Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang, Mohan Shi 외

Automatic Speech Recognition (ASR) systems struggle with child speech due to its distinct acoustic and linguistic variability and limited availability of child speech datasets, leading to high transcription error rates. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition