paper-with-me

Papers

Semi-automatic annotation of the UCU accents speech corpus

2014-05-01 · LREC 2014 5 · Rosemary Orr, Marijn Huijbregts, Roel van Beek, , Lisa Teunissen, Kate Backhouse, David van Leeuwen

Annotation and labeling of speech tasks in large multitask speech corpora is a necessary part of preparing a corpus for distribution. We address three approaches to annotation and labeling: manual, semi automatic and automatic procedures for labeling the UCU Accent Project speech data, a multilingual multitask longitudinal speech corpus. Accuracy and minimal time investment are the priorities in assessing the efficacy of each procedure. While manual labeling based on aural and visual input should produce the most accurate results, this approach is error-prone because of its repetitive nature. A semi automatic event detection system requiring manual rejection of false alarms and location and labeling of misses provided the best results. A fully automatic system could not be applied to entire speech recordings because of the variety of tasks and genres. However, it could be used to annotate separate sentences within a specific task. Acoustic confidence measures can correctly detect sentences that do not match the text with an EER of 3.3{\%}

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Event DetectionSpeech Recognition

Similar Papers 제목 키워드 기반

ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

2024-03-27 · Injy Hamed, Fadhl Eryani, David Palfreyman, Nizar Habash

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)LemmatizationPart-Of-Speech Tagging+2

English Accent Accuracy Analysis in a State-of-the-Art Automatic Speech Recognition System

2021-05-09 · Guillermo Cámbara, Alex Peiró-Lilja, Mireia Farrús, Jordi Luque

Nowadays, research in speech technologies has gotten a lot out thanks to recently created public domain corpora that contain thousands of recording hours. These large amounts of data are very helpful for training the new…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1

An Extension of the Slovak Broadcast News Corpus based on Semi-Automatic Annotation

2016-05-01 · LREC 2016 5 · Peter Viszlay, J{\'a}n Sta{\v{s}}, Tom{\'a}{\v{s}} Koct{\'u}r, Martin Lojka 외

In this paper, we introduce an extension of our previously released TUKE-BNews-SK corpus based on a semi-automatic annotation scheme. It firstly relies on the automatic transcription of the BN data performed by our Slova…

speech-recognitionSpeech Recognition

CanVEC - the Canberra Vietnamese-English Code-switching Natural Speech Corpus

2020-05-01 · LREC 2020 5 · Li Nguyen, Christopher Bryant

This paper introduces the Canberra Vietnamese-English Code-switching corpus (CanVEC), an original corpus of natural mixed speech that we semi-automatically annotated with language information, part of speech (POS) tags a…

POS

The Relevance of Text and Speech Features in Automatic Non-native English Accent Identification

2018-04-16 · Sowmya Vajjala, Ziwei Zhou

This paper describes our experiments with automatically identifying native accents from speech samples of non-native English speakers using low level audio features, and n-gram features from manual transcriptions. Using …

General ClassificationPhoneme Recognition