Learning From Failure: Data Capture in an Australian Aboriginal Community
Most low resource language technology development is premised on the need to collect data for training statistical models. When we follow the typical process of recording and transcribing text for small Indigenous languages, we hit up against the so-called “transcription bottleneck.” Therefore it is worth exploring new ways of engaging with speakers which generate data while avoiding the transcription bottleneck. We have deployed a prototype app for speakers to use for confirming system guesses in an approach to transcription based on word spotting. However, in the process of testing the app we encountered many new problems for engagement with speakers. This paper presents a close-up study of the process of deploying data capture technology on the ground in an Australian Aboriginal community. We reflect on our interactions with participants and draw lessons that apply to anyone seeking to develop methods for language data collection in an Indigenous community.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Fashioning Local Designs from Generic Speech Technologies in an Australian Aboriginal Community
An increasing number of papers have been addressing issues related to low-resource languages and the transcription bottleneck paradigm. After several years spent in Northern Australia, where some of the strongest Aborigi…
speech-recognitionSpeech RecognitionDesigning Speech Technologies for Australian Aboriginal English: Opportunities, Risks and Participation
In Australia, post-contact language varieties, including creoles and local varieties of international languages, emerged as a result of forced contact between Indigenous communities and English speakers. These contact va…
Modelling vitamin D food fortification among Aboriginal and Torres Strait Islander peoples in Australia
Background: Low vitamin D intake and high prevalence of vitamin D deficiency (serum 25-hydroxyvitamin D concentration < 50 nmol/L) among Aboriginal and Torres Strait Islander peoples highlight a need for public health st…
Multiword Expressions and the Low-Resource Scenario from the Perspective of a Local Oral Culture
Research on multiword expressions and on under-resourced languages often begins with problematisation. The existence of non-compositional meaning, or the paucity of conventional language resources, are treated as problem…
Cultural Vocal Bursts Intensity Predictionspeech-recognitionSpeech RecognitionLeveraging pre-trained representations to improve access to untranscribed speech from endangered languages
Pre-trained speech representations like wav2vec 2.0 are a powerful tool for automatic speech recognition (ASR). Yet many endangered languages lack sufficient data for pre-training such models, or are predominantly oral v…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition