Feasibility of Post-Editing Speech Transcriptions with a Mismatched Crowd
Manual correction of speech transcription can involve a selection from plausible transcriptions. Recent work has shown the feasibility of employing a mismatched crowd for speech transcription. However, it is yet to be established whether a mismatched worker has sufficiently fine-granular speech perception to choose among the phonetically proximate options that are likely to be generated from the trellis of an ASRU. Hence, we consider five languages, Arabic, German, Hindi, Russian and Spanish. For each we generate synthetic, phonetically proximate, options which emulate post-editing scenarios of varying difficulty. We consistently observe non-trivial crowd ability to choose among fine-granular options.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Clustering-based Phonetic Projection in Mismatched Crowdsourcing Channels for Low-resourced ASR
Acquiring labeled speech for low-resource languages is a difficult task in the absence of native speakers of the language. One solution to this problem involves collecting speech transcriptions from crowd workers who are…
ClusteringA case study on using speech-to-translation alignments for language documentation
For many low-resource or endangered languages, spoken language resources are more likely to be annotated with translations than with transcriptions. Recent work exploits such annotations to produce speech-to-translation …
speech-recognitionSpeech RecognitionTranslationConfidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognition
For various speech-related tasks, confidence scores from a speech recogniser are a useful measure to assess the quality of transcriptions. In traditional hidden Markov model-based automatic speech recognition (ASR) syste…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderLanguage Modeling+3A Multimodal Motion-Captured Corpus of Matched and Mismatched Extravert-Introvert Conversational Pairs
This paper presents a new corpus, the Personality Dyads Corpus, consisting of multimodal data for three conversations between three personality-matched, two-person dyads (a total of 9 separate dialogues). Participants we…
Grammatical error detection in transcriptions of spoken English
We describe the collection of transcription corrections and grammatical error annotations for the CrowdED Corpus of spoken English monologues on business topics. The corpus recordings were crowdsourced from native speake…
Grammatical Error CorrectionGrammatical Error Detection