paper-with-me

홈 › Papers

Enhancing Documentation of Hupa with Automatic Speech Recognition

2022-05-01 · ComputEL (ACL) 2022 5 · Zoey Liu, Justin Spence, Emily Tucker Prud’hommeaux

This study investigates applications of automatic speech recognition (ASR) techniques to Hupa, a critically endangered Native American language from the Dene (Athabaskan) language family. Using around 9h12m of spoken data produced by one elder who is a first-language Hupa speaker, we experimented with different evaluation schemes and training settings. On average a fully connected deep neural network reached a word error rate of 35.26%. Our overall results illustrate the utility of ASR for making Hupa language documentation more accessible and usable. In addition, we found that when training acoustic models, using recordings with transcripts that were not carefully verified did not necessarily have a negative effect on model performance. This shows promise for speech corpora of indigenous languages that commonly include transcriptions produced by second-language speakers or linguists who have advanced knowledge in the language of interest.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Prabhupadavani: A Code-mixed Speech Translation Data for 25 Languages

2022-01-27 · LaTeCHCLfL (COLING) 2022 10 · Jivnesh Sandhan, Ayush Daksh, Om Adideva Paranjay, Laxmidhar Behera 외

Nowadays, the interest in code-mixing has become ubiquitous in Natural Language Processing (NLP); however, not much attention has been given to address this phenomenon for Speech Translation (ST) task. This can be solely…

Cultural Vocal Bursts Intensity PredictionMachine TranslationTranslation

Automatic Documentation of ICD Codes with Far-Field Speech Recognition

2018-04-30 · Albert Haque, Corinna Fukushima

Documentation errors increase healthcare costs and cause unnecessary patient deaths. As the standard language for diagnoses and billing, ICD codes serve as the foundation for medical documentation worldwide. Despite the …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Highland Puebla Nahuatl Speech Translation Corpus for Endangered Language Documentation

2021-06-01 · NAACL (AmericasNLP) 2021 6 · Jiatong Shi, Jonathan D. Amith, Xuankai Chang, Siddharth Dalmia 외

Documentation of endangered languages (ELs) has become increasingly urgent as thousands of languages are on the verge of disappearing by the end of the 21st century. One challenging aspect of documentation is to develop …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BIG-bench Machine LearningMachine Translation+3

Supporting SENCOTEN Language Documentation Efforts with Automatic Speech Recognition

2025-07-14 · Mengzhe Geng, Patrick Littell, Aidan Pine, PENÁĆ 외 arxiv

The SENCOTEN language, spoken on the Saanich peninsula of southern Vancouver Island, is in the midst of vigorous language revitalization efforts to turn the tide of language loss as a result of colonial language policies…

Cross-Lingual TransferSpeech Recognition

Towards Building an Automatic Transcription System for Language Documentation: Experiences from Muyu

2020-05-01 · LREC 2020 5 · Alex Zahrer, er, Andrej Zgank, Barbara Schuppler

Since at least half of the world{'}s 6000 plus languages will vanish during the 21st century, language documentation has become a rapidly growing field in linguistics. A fundamental challenge for language documentation i…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phoneme Recognitionspeech-recognition+1