paper-with-me

홈 › Papers

Dataset of British English speech recordings for psychoacoustics and speech processing research

2022-02-15 · Data in Brief 2022 2 · Trevor John Cox, Simone Graetzer, Michael A Akeroyd, Jonathan Barker, John Culling, Graham Naylor, Eszter Porter, Rhoddy Viveros Muñoz

The Clarity Speech Corpus is a forty speaker British English speech dataset. The corpus was created for the purpose of running listening tests to gauge speech intelligibility and quality in the Clarity Project, which has the goal of advancing speech signal processing by hearing aids through a series of challenges. The dataset is suitable for machine learning and other uses in speech and hearing technology, acoustics and psychoacoustics. The data comprises recordings of approximately 10,000 sentences drawn from the British National Corpus (BNC) with suitable length, words and grammatical construction for speech intelligibility testing. The collection process involved the selection of a subset of BNC sentences, the recording of these produced by 40 British English speakers, and the processing of these recordings to create individual sentence recordings with associated prompts and metadata.

📄 PDF Abstract BibTeX

Code (1)

claritychallenge/clarity 공식 구현

Tasks

Sentence

Similar Papers 제목 키워드 기반

Open-source Multi-speaker Corpora of the English Accents in the British Isles

2020-05-01 · LREC 2020 5 · Isin Demirsahin, Oddur Kjartansson, Alex Gutkin, er 외

This paper presents a dataset of transcribed high-quality audio of English sentences recorded by volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as…

A Framework for Collecting Realistic Recordings of Dysarthric Speech - the homeService Corpus

2016-05-01 · LREC 2016 5 · Mauro Nicolao, Heidi Christensen, Stuart Cunningham, Phil Green 외

This paper introduces a new British English speech database, named the homeService corpus, which has been gathered as part of the homeService project. This project aims to help users with speech and motor disabilities to…

Identifications of Speaker Ethnicity in South-East England: Multicultural London English as a Divisible Perceptual Variety

2020-05-01 · LREC 2020 5 · Am Cole, a

This study uses crowdsourcing through LanguageARC to collect data on levels of accuracy in the identification of speakers{'} ethnicities. Ten participants (5 US; 5 South-East England) classified lexically identical speec…

A Corpus of Spontaneous Multi-party Conversation in Bosnian Serbo-Croatian and British English

2012-05-01 · LREC 2012 5 · Emina Kurti{\'c}, Bill Wells, Guy J. Brown, Timothy Kempton 외

In this paper we present a corpus of audio and video recordings of spontaneous, face-to-face multi-party conversation in two languages. Freely available high quality recordings of mundane, non-institutional, multi-party …

SNuC: The Sheffield Numbers Spoken Language Corpus

2022-06-01 · LREC 2022 6 · Emma Barker, Jon Barker, Robert Gaizauskas, Ning Ma 외

We present SNuC, the first published corpus of spoken alphanumeric identifiers of the sort typically used as serial and part numbers in the manufacturing sector. The dataset contains recordings and transcriptions of over…