Dataset of British English speech recordings for psychoacoustics and speech processing research
The Clarity Speech Corpus is a forty speaker British English speech dataset. The corpus was created for the purpose of running listening tests to gauge speech intelligibility and quality in the Clarity Project, which has the goal of advancing speech signal processing by hearing aids through a series of challenges. The dataset is suitable for machine learning and other uses in speech and hearing technology, acoustics and psychoacoustics. The data comprises recordings of approximately 10,000 sentences drawn from the British National Corpus (BNC) with suitable length, words and grammatical construction for speech intelligibility testing. The collection process involved the selection of a subset of BNC sentences, the recording of these produced by 40 British English speakers, and the processing of these recordings to create individual sentence recordings with associated prompts and metadata.
Code (1)
Tasks
SentenceSimilar Papers 제목 키워드 기반
Open-source Multi-speaker Corpora of the English Accents in the British Isles
This paper presents a dataset of transcribed high-quality audio of English sentences recorded by volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as…
A Framework for Collecting Realistic Recordings of Dysarthric Speech - the homeService Corpus
This paper introduces a new British English speech database, named the homeService corpus, which has been gathered as part of the homeService project. This project aims to help users with speech and motor disabilities to…
Identifications of Speaker Ethnicity in South-East England: Multicultural London English as a Divisible Perceptual Variety
This study uses crowdsourcing through LanguageARC to collect data on levels of accuracy in the identification of speakers{'} ethnicities. Ten participants (5 US; 5 South-East England) classified lexically identical speec…
A Corpus of Spontaneous Multi-party Conversation in Bosnian Serbo-Croatian and British English
In this paper we present a corpus of audio and video recordings of spontaneous, face-to-face multi-party conversation in two languages. Freely available high quality recordings of mundane, non-institutional, multi-party …
SNuC: The Sheffield Numbers Spoken Language Corpus
We present SNuC, the first published corpus of spoken alphanumeric identifiers of the sort typically used as serial and part numbers in the manufacturing sector. The dataset contains recordings and transcriptions of over…