paper-with-me

Papers

Palabras: Crowdsourcing Transcriptions of L2 Speech

2016-05-01 · LREC 2016 5 · S, Eric ers, Pepi Burgos, Catia Cucchiarini, Roel van Hout,

We developed a web application for crowdsourcing transcriptions of Dutch words spoken by Spanish L2 learners. In this paper we discuss the design of the application and the influence of metadata and various forms of feedback. Useful data were obtained from 159 participants, with an average of over 20 transcriptions per item, which seems a satisfactory result for this type of research. Informing participants about how many items they still had to complete, and not how many they had already completed, turned to be an incentive to do more items. Assigning participants a score for their performance made it more attractive for them to carry out the transcription task, but this seemed to influence their performance. We discuss possible advantages and disadvantages in connection with the aim of the research and consider possible lessons for designing future experiments.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Clustering-based Phonetic Projection in Mismatched Crowdsourcing Channels for Low-resourced ASR

2016-12-01 · WS 2016 12 · Wenda Chen, Mark Hasegawa-Johnson, Nancy Chen, Preethi Jyothi 외

Acquiring labeled speech for low-resource languages is a difficult task in the absence of native speakers of the language. One solution to this problem involves collecting speech transcriptions from crowd workers who are…

Clustering

CrowdSpeech and VoxDIY: Benchmark Datasets for Crowdsourced Audio Transcription

2021-07-02 · Nikita Pavlichenko, Ivan Stelmakh, Dmitry Ustalov

Domain-specific data is the crux of the successful transfer of machine learning systems from benchmarks to real life. In simple problems such as image classification, crowdsourcing has become one of the standard tools fo…

Crowdsourced Text Aggregationimage-classificationspeech-recognition

Free English and Czech telephone speech corpus shared under the CC-BY-SA 3.0 license

2014-05-01 · LREC 2014 5 · Mat{\v{e}}j Korvas, Ond{\v{r}}ej Pl{\'a}tek, Ond{\v{r}}ej Du{\v{s}}ek, Luk{\'a}{\v{s}} {\v{Z}}ilka 외

We present a dataset of telephone conversations in English and Czech, developed for training acoustic models for automatic speech recognition (ASR) in spoken dialogue systems (SDSs). The data comprise 45 hours of speech …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognition+2

A Systematic Approach to Derive a Refined Speech Corpus for Sinhala

2022-06-01 · LREC 2022 6 · Disura Warusawithana, Nilmani Kulaweera, Lakshan Weerasinghe, Buddhika Karunarathne

Speech Recognition is an active research area where advances of technology have continuously driven the development of research work. However, due to the lack of adequate resources, certain languages such as Sinhala, are…

speech-recognitionSpeech Recognition

A case study on using speech-to-translation alignments for language documentation

2017-02-14 · WS 2017 3 · Antonios Anastasopoulos, David Chiang

For many low-resource or endangered languages, spoken language resources are more likely to be annotated with translations than with transcriptions. Recent work exploits such annotations to produce speech-to-translation …

speech-recognitionSpeech RecognitionTranslation