paper-with-me

Papers

Back-Translation-Style Data Augmentation for Mandarin Chinese Polyphone Disambiguation

2022-11-17 · Chunyu Qiang, Peng Yang, Hao Che, Jinba Xiao, Xiaorui Wang, Zhongyuan Wang

Conversion of Chinese Grapheme-to-Phoneme (G2P) plays an important role in Mandarin Chinese Text-To-Speech (TTS) systems, where one of the biggest challenges is the task of polyphone disambiguation. Most of the previous polyphone disambiguation models are trained on manually annotated datasets, and publicly available datasets for polyphone disambiguation are scarce. In this paper we propose a simple back-translation-style data augmentation method for mandarin Chinese polyphone disambiguation, utilizing a large amount of unlabeled text data. Inspired by the back-translation technique proposed in the field of machine translation, we build a Grapheme-to-Phoneme (G2P) model to predict the pronunciation of polyphonic character, and a Phoneme-to-Grapheme (P2G) model to predict pronunciation into text. Meanwhile, a window-based matching strategy and a multi-model scoring strategy are proposed to judge the correctness of the pseudo-label. We design a data balance strategy to improve the accuracy of some typical polyphonic characters in the training set with imbalanced distribution or data scarcity. The experimental result shows the effectiveness of the proposed back-translation-style data augmentation method.

📄 PDF Abstract BibTeX arXiv:2211.09495

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationMachine TranslationPolyphone disambiguationPseudo Labeltext-to-speechText to SpeechTranslation

Similar Papers 제목 키워드 기반

Text Style Transfer Back-Translation

2023-06-02 · Daimeng Wei, Zhanglin Wu, Hengchao Shang, Zongyao Li 외

Back Translation (BT) is widely used in the field of machine translation, as it has been proved effective for enhancing translation quality. However, BT mainly improves the translation of inputs that share a similar styl…

Data AugmentationDomain AdaptationMachine TranslationStyle Transfer+2

Enhancing Taiwanese Hokkien Dual Translation by Exploring and Standardizing of Four Writing Systems

2024-03-18 · Bo-Han Lu, Yi-Hsuan Lin, En-Shiun Annie Lee, Richard Tzong-Han Tsai

Machine translation focuses mainly on high-resource languages (HRLs), while low-resource languages (LRLs) like Taiwanese Hokkien are relatively under-explored. The study aims to address this gap by developing a dual tran…

Machine TranslationTranslation

Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech

2024-03-25 · Shannon Wotherspoon, William Hartmann, Matthew Snover

This paper introduces a set of English translations for a 123-hour subset of the CallHome Mandarin Chinese data and the HKUST Mandarin Telephone Speech data for the task of speech translation. Paired source-language spee…

Translation

The Xiaomi Text-to-Text Simultaneous Speech Translation System for IWSLT 2022

2022-05-01 · IWSLT (ACL) 2022 5 · Bao Guo, Mengge Liu, Wen Zhang, Hexuan Chen 외

This system paper describes the Xiaomi Translation System for the IWSLT 2022 Simultaneous Speech Translation (noted as SST) shared task. We participate in the English-to-Mandarin Chinese Text-to-Text (noted as T2T) track…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationKnowledge Distillation+4

FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation

2022-10-01 · Parker Riley, Timothy Dozat, Jan A. Botha, Xavier Garcia 외

We present FRMT, a new dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation. The dataset consists of professional translations from English into two regiona…

Machine TranslationTranslation