paper-with-me

Papers

A Mandarin-English Code-Switching Corpus

2012-05-01 · LREC 2012 5 · Ying Li, Yue Yu, Pascale Fung

Generally the existing monolingual corpora are not suitable for large vocabulary continuous speech recognition (LVCSR) of code-switching speech. The motivation of this paper is to study the rules and constraints code-switching follows and design a corpus for code-switching LVCSR task. This paper presents the development of a Mandarin-English code-switching corpus. This corpus consists of four parts: 1) conversational meeting speech and its data; 2) project meeting speech data; 3) student interviews speech; 4) text data of on-line news. The speech was transcribed by an annotator and verified by Mandarin-English bilingual speakers manually. We propose an approach for automatically downloading from the web text data that contains code-switching. The corpus includes both intra-sentential code-switching (switch in the middle of a sentence) and inter-sentential code-switching (switch at the end of the sentence). The distribution of part-of-speech (POS) tags and code-switching reasons are reported.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Boundary DetectionLanguage IdentificationPOSSentencespeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

TALCS: An Open-Source Mandarin-English Code-Switching Corpus and a Speech Recognition Baseline

2022-06-27 · Chengfei Li, Shuhao Deng, Yaoping Wang, Guangjing Wang 외

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real on…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

On the End-to-End Solution to Mandarin-English Code-switching Speech Recognition

2018-11-01 · Zhiping Zeng, Yerbolat Khassanov, Van Tung Pham, Hai-Hua Xu 외

Code-switching (CS) refers to a linguistic phenomenon where a speaker uses different languages in an utterance or between alternating utterances. In this work, we study end-to-end (E2E) approaches to the Mandarin-English…

Data AugmentationLanguage IdentificationLanguage ModelingLanguage Modelling+2

Non-autoregressive Mandarin-English Code-switching Speech Recognition

2021-04-06 · Shun-Po Chuang, Heng-Jui Chang, Sung-Feng Huang, Hung-Yi Lee

Mandarin-English code-switching (CS) is frequently used among East and Southeast Asian people. However, the intra-sentence language switching of the two very different languages makes recognizing CS speech challenging. M…

DecoderSentencespeech-recognitionSpeech Recognition

Introducing MELI: the Mandarin-English Language Interview Corpus

2026-03-27 · Suyuan Liu, Molly Babel arxiv

We introduce the Mandarin-English Language Interview (MELI) Corpus, an open-source resource of 29.8 hours of speech from 51 Mandarin-English bilingual speakers. MELI combines matched sessions in Mandarin and English with…

Simple yet Effective Code-Switching Language Identification with Multitask Pre-Training and Transfer Learning

2023-05-31 · Shuyue Stella Li, Cihan Xiao, Tianjian Li, Bismarck Odoom

Code-switching, also called code-mixing, is the linguistics phenomenon where in casual settings, multilingual speakers mix words from different languages in one utterance. Due to its spontaneous nature, code-switching is…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Identification+3