paper-with-me

홈 › Papers

Fast transcription of speech in low-resource languages

2019-09-16 · Mark Hasegawa-Johnson, Camille Goudeseune, Gina-Anne Levow

We present software that, in only a few hours, transcribes forty hours of recorded speech in a surprise language, using only a few tens of megabytes of noisy text in that language, and a zero-resource grapheme to phoneme (G2P) table. A pretrained acoustic model maps acoustic features to phonemes; a reversed G2P maps these to graphemes; then a language model maps these to a most-likely grapheme sequence, i.e., a transcription. This software has worked successfully with corpora in Arabic, Assam, Kinyarwanda, Russian, Sinhalese, Swahili, Tagalog, and Tamil.

📄 PDF Abstract BibTeX arXiv:1909.07285

Code (1)

uiuc-sst/asr24 공식 구현

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Endangered Language Documentation: Bootstrapping a Chatino Speech Corpus, Forced Aligner, ASR

2016-05-01 · LREC 2016 5 · Malgorzata {\'C}avar, Damir {\'C}avar, Hilaria Cruz

This project approaches the problem of language documentation and revitalization from a rather untraditional angle. To improve and facilitate language documentation of endangered languages, we attempt to use corpus lingu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

An Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource Languages

2016-09-26 · EMNLP 2016 11 · Antonios Anastasopoulos, David Chiang, Long Duong

For many low-resource languages, spoken language resources are more likely to be annotated with translations than with transcriptions. Translated speech data is potentially valuable for documenting endangered languages o…

ClusteringDynamic Time WarpingTranslation

Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning

2026-07-07 · Dylan Luke Holyoak arxiv

Leading multilingual speech recognition models like Whisper transcribe diverse, low-resource languages without language-specific training but are computationally expensive to deploy. Token merging mitigates this ineffici…

Cost Analysis of Human-corrected Transcription for Predominately Oral Languages

2025-10-14 · Yacouba Diarra, Nouhoum Souleymane Coulibaly, Michael Leventhal arxiv

Creating speech datasets for low-resource languages is a critical yet poorly understood challenge, particularly regarding the actual cost in human labor. This paper investigates the time and complexity required to produc…

Phonemic Representation and Transcription for Speech to Text Applications for Under-resourced Indigenous African Languages: The Case of Kiswahili

2022-10-29 · Ebbie Awino, Lilian Wanzare, Lawrence Muchemi, Barack Wanjawa 외

Building automatic speech recognition (ASR) systems is a challenging task, especially for under-resourced languages that need to construct corpora nearly from scratch and lack sufficient training data. It has emerged tha…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1