paper-with-me

홈 › Papers

Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech

2024-03-25 · Shannon Wotherspoon, William Hartmann, Matthew Snover

This paper introduces a set of English translations for a 123-hour subset of the CallHome Mandarin Chinese data and the HKUST Mandarin Telephone Speech data for the task of speech translation. Paired source-language speech and target-language text is essential for training end-to-end speech translation systems and can provide substantial performance improvements for cascaded systems as well, relative to training on more widely available text data sets. We demonstrate that fine-tuning a general-purpose translation model to our Mandarin-English conversational telephone speech training set improves target-domain BLEU by more than 8 points, highlighting the importance of matched training data.

📄 PDF Abstract BibTeX arXiv:2404.11619

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

BSTC: A Large-Scale Chinese-English Speech Translation Dataset

2021-04-08 · NAACL (AutoSimTrans) 2021 6 · Ruiqing Zhang, Xiyang Wang, Chuanqiang Zhang, Zhongjun He 외

This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

TALCS: An Open-Source Mandarin-English Code-Switching Corpus and a Speech Recognition Baseline

2022-06-27 · Chengfei Li, Shuhao Deng, Yaoping Wang, Guangjing Wang 외

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real on…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Mandarin-English Code-Switching Corpus

2012-05-01 · LREC 2012 5 · Ying Li, Yue Yu, Pascale Fung

Generally the existing monolingual corpora are not suitable for large vocabulary continuous speech recognition (LVCSR) of code-switching speech. The motivation of this paper is to study the rules and constraints code-swi…

Boundary DetectionLanguage IdentificationPOSSentence+2

Detecting dementia in Mandarin Chinese using transfer learning from a parallel corpus

2019-03-03 · NAACL 2019 6 · Bai Li, Yi-Te Hsu, Frank Rudzicz

Machine learning has shown promise for automatic detection of Alzheimer's disease (AD) through speech; however, efforts are hampered by a scarcity of data, especially in languages other than English. We propose a method …

BIG-bench Machine LearningMachine TranslationTransfer LearningTranslation

The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results

2020-07-12 · Xian Shi, Qiangze Feng, Lei Xie

Code-switching (CS) is a common phenomenon and recognizing CS speech is challenging. But CS speech data is scarce and there' s no common testbed in relevant research. This paper describes the design and main outcomes of …

Data AugmentationLanguage Identificationspeech-recognitionSpeech Recognition