paper-with-me

홈 › Papers

Phoneme transcription of endangered languages: an evaluation of recent ASR architectures in the single speaker scenario

2022-05-01 · Findings (ACL) 2022 5 · Gilles Boulianne

Transcription is often reported as the bottleneck in endangered language documentation, requiring large efforts from scarce speakers and transcribers. In general, automatic speech recognition (ASR) can be accurate enough to accelerate transcription only if trained on large amounts of transcribed data. However, when a single speaker is involved, several studies have reported encouraging results for phonetic transcription even with small amounts of training. Here we expand this body of work on speaker-dependent transcription by comparing four ASR approaches, notably recent transformer and pretrained multilingual models, on a common dataset of 11 languages. To automate data preparation, training and evaluation steps, we also developed a phoneme recognition setup which handles morphologically complex languages and writing systems for which no pronunciation dictionary exists.We find that fine-tuning a multilingual pretrained model yields an average phoneme error rate (PER) of 15% for 6 languages with 99 minutes or less of transcribed data for training. For the 5 languages with between 100 and 192 minutes of training, we achieved a PER of 8.4% or less. These results on a number of varied languages suggest that ASR can now significantly reduce transcription efforts in the speaker-dependent situation common in endangered language work.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phoneme Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Universal Automatic Phonetic Transcription into the International Phonetic Alphabet

2023-08-07 · Chihiro Taguchi, Yusuke Sakai, Parisa Haghani, David Chiang

This paper presents a state-of-the-art model for transcribing speech in any language into the International Phonetic Alphabet (IPA). Transcription of spoken languages into IPA is an essential yet time-consuming process i…

AlloVera: A Multilingual Allophone Database

2020-04-17 · LREC 2020 5 · David R. Mortensen, Xinjian Li, Patrick Littell, Alexis Michaud 외

We introduce a new resource, AlloVera, which provides mappings from 218 allophones to phonemes for 14 languages. Phonemes are contrastive phonological units, and allophones are their various concrete realizations, which …

speech-recognitionSpeech Recognition

Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages

2026-04-20 · V. S. D. S. Mahesh Akavarapu, Michael Daniel, Gerhard Jäger arxiv

We present a phoneme-level analysis of automatic speech recognition (ASR) for two low-resourced and phonologically complex East Caucasian languages, Archi and Rutul, based on curated and standardized speech-transcript re…

Speech Recognition

Endangered Language Documentation: Bootstrapping a Chatino Speech Corpus, Forced Aligner, ASR

2016-05-01 · LREC 2016 5 · Malgorzata {\'C}avar, Damir {\'C}avar, Hilaria Cruz

This project approaches the problem of language documentation and revitalization from a rather untraditional angle. To improve and facilitate language documentation of endangered languages, we attempt to use corpus lingu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

User-Centric Evaluation of OCR Systems for Kwak'wala

2023-02-26 · Shruti Rijhwani, Daisy Rosenblum, Michayla King, Antonios Anastasopoulos 외

There has been recent interest in improving optical character recognition (OCR) for endangered languages, particularly because a large number of documents and books in these languages are not in machine-readable formats.…

Optical Character RecognitionOptical Character Recognition (OCR)