paper-with-me

Papers

Exploiting Low-Resource Code-Switching Data to Mandarin-English Speech Recognition Systems

2021-10-01 · ROCLING 2021 10 · Hou-An Lin, Chia-Ping Chen

In this paper, we investigate how to use limited code-switching data to implement a code-switching speech recognition system. We utilize the Transformer end-to-end model to develop our code switching speech recognition system, which is trained with the Mandarin dataset and a small amount of Mandarin-English code switching dataset, as the baseline of this paper. Next, we compare the performance of systems after adding multi-task learning and transfer learning. Character Error Rate(CER) is adopted as the criterion for the system. Finally, we combined the three systems with the language model, respectively, our best result dropped to 23.9% compared with the baseline of 28.7%.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMulti-Task Learningspeech-recognitionSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Simple yet Effective Code-Switching Language Identification with Multitask Pre-Training and Transfer Learning

2023-05-31 · Shuyue Stella Li, Cihan Xiao, Tianjian Li, Bismarck Odoom

Code-switching, also called code-mixing, is the linguistics phenomenon where in casual settings, multilingual speakers mix words from different languages in one utterance. Due to its spontaneous nature, code-switching is…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Identification+3

A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario

2024-12-01 · Zheshu Song, Ziyang Ma, Yifan Yang, Jianheng Zhuo 외

Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in the Automatic Speech Recognition (ASR) fi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Monolingual Data Selection Analysis for English-Mandarin Hybrid Code-switching Speech Recognition

2020-09-14

In this paper, we conduct data selection analysis in building an English-Mandarin code-switching (CS) speech recognition (CSSR) system, which is aimed for a real CSSR contest in China. The overall training sets have thre…

speech-recognitionSpeech Recognition

A Mandarin-English Code-Switching Corpus

2012-05-01 · LREC 2012 5 · Ying Li, Yue Yu, Pascale Fung

Generally the existing monolingual corpora are not suitable for large vocabulary continuous speech recognition (LVCSR) of code-switching speech. The motivation of this paper is to study the rules and constraints code-swi…

Boundary DetectionLanguage IdentificationPOSSentence+2

The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results

2020-07-12 · Xian Shi, Qiangze Feng, Lei Xie

Code-switching (CS) is a common phenomenon and recognizing CS speech is challenging. But CS speech data is scarce and there' s no common testbed in relevant research. This paper describes the design and main outcomes of …

Data AugmentationLanguage Identificationspeech-recognitionSpeech Recognition