paper-with-me

Papers

ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn Conversation

2021-12-12 · LREC 2022 6 · Holy Lovenia, Samuel Cahyawijaya, Genta Indra Winata, Peng Xu, Xu Yan, Zihan Liu, Rita Frieske, Tiezheng Yu, Wenliang Dai, Elham J. Barezi, Qifeng Chen, Xiaojuan Ma, Bertram E. Shi, Pascale Fung

Code-switching is a speech phenomenon occurring when a speaker switches language during a conversation. Despite the spontaneous nature of code-switching in conversational spoken language, most existing works collect code-switching data from read speech instead of spontaneous speech. ASCEND (A Spontaneous Chinese-English Dataset) is a high-quality Mandarin Chinese-English code-switching corpus built on spontaneous multi-turn conversational dialogue sources collected in Hong Kong. We report ASCEND's design and procedure for collecting the speech data, including annotations. ASCEND consists of 10.62 hours of clean speech, collected from 23 bilingual speakers of Chinese and English. Furthermore, we conduct baseline experiments using pre-trained wav2vec 2.0 models, achieving a best performance of 22.69\% character error rate and 27.05% mixed error rate.

📄 PDF Abstract BibTeX arXiv:2112.06223

Code (2)

HLTCHKUST/ASCEND 공식 구현 pytorch
jasonppy/promptingwhisper pytorch

Similar Papers 제목 키워드 기반

Code-switching in text and speech reveals information-theoretic audience design

2024-08-08 · Debasmita Bhattacharya, Marten Van Schijndel

In this work, we use language modeling to investigate the factors that influence code-switching. Code-switching occurs when a speaker alternates between one language variety (the primary language) and another (the second…

Language ModelingLanguage Modelling

Optimizing Bilingual Neural Transducer with Synthetic Code-switching Text Generation

2022-10-21 · Thien Nguyen, Nathalie Tran, Liuhui Deng, Thiago Fraga da Silva 외

Code-switching describes the practice of using more than one language in the same sentence. In this study, we investigate how to optimize a neural transducer based bilingual automatic speech recognition (ASR) model for c…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+2

PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions

2026-05-18 · Sicheng Jin, Dipankar Srirag, Aditya Joshi arxiv

While modern Automatic Speech Recognition (ASR) systems achieve high accuracy on benchmark corpora, their performance often degrades when there is real-world variability. This work focuses on variability arising due to a…

Speech Recognition

Automatic Speech Recognition for Non-Native English: Accuracy and Disfluency Handling

2025-03-10 · Michael McGuire

Automatic speech recognition (ASR) has been an essential component of computer assisted language learning (CALL) and computer assisted language testing (CALT) for many years. As this technology continues to develop rapid…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

ChiEngMixBench: Evaluating Large Language Models on Spontaneous and Natural Chinese-English Code-Mixed Generation

2026-01-02 · Qingyan Yang, Tongxi Wang, Yunsheng Luo arxiv

Code-mixing is increasingly prevalent in interactions between humans and large language models, yet existing work often reduces it to a translation or convertibility problem, making it difficult to assess whether a model…