paper-with-me

홈 › Papers

Frustratingly Easy Data Augmentation for Low-Resource ASR

2025-09-18 · Katsumi Ibaraki, David Chiang arxiv

This paper introduces three self-contained data augmentation methods for low-resource Automatic Speech Recognition (ASR). Our techniques first generate novel text--using gloss-based replacement, random replacement, or an LLM-based approach--and then apply Text-to-Speech (TTS) to produce synthetic audio. We apply these methods, which leverage only the original annotated data, to four languages with extremely limited resources (Vatlongos, Nashta, Shinekhen Buryat, and Kakabe). Fine-tuning a pretrained Wav2Vec2-XLSR-53 model on a combination of the original audio and generated synthetic data yields significant performance gains, including a 14.3% absolute WER reduction for Nashta. The methods prove effective across all four low-resource languages and also show utility for high-resource languages like English, demonstrating their broad applicability.

📄 PDF Abstract BibTeX arXiv:2509.15373

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionData Augmentation

Similar Papers 제목 키워드 기반

Frustratingly Easy Uncertainty Estimation for Distribution Shift

2021-06-07 · Tiago Salvador, Vikram Voleti, Alexander Iannantuono, Adam Oberman

Distribution shift is an important concern in deep image classification, produced either by corruption of the source images, or a complete change, with the solution involving domain adaptation. While the primary goal is …

Domain Adaptationimage-classificationImage ClassificationUnsupervised Domain Adaptation

Frustratingly Easy Cross-Lingual Transfer for Transition-Based Dependency Parsing

2016-06-01 · NAACL 2016 6 · Oph{\'e}lie Lacroix, Lauriane Aufrant, Guillaume Wisniewski, Fran{\c{c}}ois Yvon
Cross-Lingual TransferDependency ParsingTransition-Based Dependency Parsing

TripleE: Easy Domain Generalization via Episodic Replay

2022-10-04 · Xiaomeng Li, Hongyu Ren, Huifeng Yao, Ziwei Liu

Learning how to generalize the model to unseen domains is an important area of research. In this paper, we propose TripleE, and the main idea is to encourage the network to focus on training on subsets (learning with rep…

DiversityDomain Generalization

Frustratingly Easy Neural Domain Adaptation

2016-12-01 · COLING 2016 12 · Young-Bum Kim, Karl Stratos, Ruhi Sarikaya

Popular techniques for domain adaptation such as the feature augmentation method of Daum{\'e} III (2009) have mostly been considered for sparse binary-valued features, but not for dense real-valued features such as those…

Domain Adaptation

Frustratingly Easy Natural Question Answering

2019-09-11 · Lin Pan, Rishav Chakravarti, Anthony Ferritto, Michael Glass 외

Existing literature on Question Answering (QA) mostly focuses on algorithmic novelty, data augmentation, or increasingly large pre-trained language models like XLNet and RoBERTa. Additionally, a lot of systems on the QA …

Data AugmentationNatural QuestionsQuestion AnsweringTransfer Learning