paper-with-me

홈 › Papers

SNuC: The Sheffield Numbers Spoken Language Corpus

2022-06-01 · LREC 2022 6 · Emma Barker, Jon Barker, Robert Gaizauskas, Ning Ma, Monica Lestari Paramita

We present SNuC, the first published corpus of spoken alphanumeric identifiers of the sort typically used as serial and part numbers in the manufacturing sector. The dataset contains recordings and transcriptions of over 50 native British English speakers, speaking over 13,000 multi-character alphanumeric sequences and totalling almost 20 hours of recorded speech. We describe requirements taken into account in the designing the corpus and the methodology used to construct it. We present summary statistics describing the corpus contents, as well as a preliminary investigation into errors in spoken alphanumeric identifiers. We validate the corpus by showing how it can be used to adapt a deep learning neural network based ASR system, resulting in improved recognition accuracy on the task of spoken alphanumeric identifier recognition. Finally, we discuss further potential uses for the corpus and for the tools developed to construct it.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Numbers Normalisation in the Inflected Languages: a Case Study of Polish

2019-08-01 · WS 2019 8 · Rafa{\l} Po{\'s}wiata, Micha{\l} Pere{\l}kiewicz

Text normalisation in Text-to-Speech systems is a process of converting written expressions to their spoken forms. This task is complicated because in many cases the normalised form depends on the context. Furthermore, w…

text-to-speechText to Speech

The USFD Spoken Language Translation System for IWSLT 2014

2015-09-13 · Raymond W. M. Ng, Mortaza Doulaty, Rama Doddipatla, Wilker Aziz 외

The University of Sheffield (USFD) participated in the International Workshop for Spoken Language Translation (IWSLT) in 2014. In this paper, we will introduce the USFD SLT system for IWSLT. Automatic speech recognition …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+4

Designing a Speech Corpus for the Development and Evaluation of Dictation Systems in Latvian

2016-05-01 · LREC 2016 5 · M{\=a}rcis Pinnis, Askars Salimbajevs, Ilze Auzi{\c{n}}a

In this paper the authors present a speech corpus designed and created for the development and evaluation of dictation systems in Latvian. The corpus consists of over nine hours of orthographically annotated speech from …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+3

Timers and Such: A Practical Benchmark for Spoken Language Understanding with Numbers

2021-04-04 · Loren Lugosch, Piyush Papreja, Mirco Ravanelli, Abdelwahab Heba 외

This paper introduces Timers and Such, a new open source dataset of spoken English commands for common voice control use cases involving numbers. We describe the gap in existing spoken language understanding datasets tha…

Spoken Language Understanding

Sheffield Submissions for the WMT18 Quality Estimation Shared Task

2018-10-01 · WS 2018 10 · Julia Ive, Carolina Scarton, Fr{\'e}d{\'e}ric Blain, Lucia Specia

In this paper we present the University of Sheffield submissions for the WMT18 Quality Estimation shared task. We discuss our submissions to all four sub-tasks, where ours is the only team to participate in all language …

AllMachine Translation