paper-with-me

Papers

Data augmentation for low-resource grapheme-to-phoneme mapping

2021-08-01 · ACL (SIGMORPHON) 2021 8 · Michael Hammond

In this paper we explore a very simple neural approach to mapping orthography to phonetic transcription in a low-resource context. The basic idea is to start from a baseline system and focus all efforts on data augmentation. We will see that some techniques work, but others do not.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

2026-05-01 · Priyam Mazumdar, Yurii Halychanskyi, Steven Guo, Mark Hasegawa-Johnson 외 arxiv

Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P)…

Speech Synthesis

Mining Large-Scale Low-Resource Pronunciation Data From Wikipedia

2021-01-27 · Tania Chakraborty, Manasa Prasad, Theresa Breiner, Sandy Ritchie 외

Pronunciation modeling is a key task for building speech technology in new languages, and while solid grapheme-to-phoneme (G2P) mapping systems exist, language coverage can stand to be improved. The information needed to…

Ensemble Self-Training for Low-Resource Languages: Grapheme-to-Phoneme Conversion and Morphological Inflection

2020-07-01 · WS 2020 7 · Xiang Yu, Ngoc Thang Vu, Jonas Kuhn

We present an iterative data augmentation framework, which trains and searches for an optimal ensemble and simultaneously annotates new training data in a self-training style. We apply this framework on two SIGMORPHON 20…

Data AugmentationGrapheme-to-Phoneme ConversionMorphological Inflection

Grapheme-Based Cross-Language Forced Alignment: Results with Uralic Languages

2021-05-01 · NoDaLiDa 2021 5 · Juho Leinonen, Sami Virpioja, Mikko Kurimo

Forced alignment is an effective process to speed up linguistic research. However, most forced aligners are language-dependent, and under-resourced languages rarely have enough resources to train an acoustic model for an…

A systematic comparison of grapheme-based vs. phoneme-based label units for encoder-decoder-attention models

2020-05-19 · Mohammad Zeineldeen, Albert Zeyer, Wei Zhou, Thomas Ng 외

Following the rationale of end-to-end modeling, CTC, RNN-T or encoder-decoder-attention models for automatic speech recognition (ASR) use graphemes or grapheme-based subword units based on e.g. byte-pair encoding (BPE). …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1