paper-with-me

Papers

BASPRO: a balanced script producer for speech corpus collection based on the genetic algorithm

2022-12-11 · Yu-Wen Chen, Hsin-Min Wang, Yu Tsao

The performance of speech-processing models is heavily influenced by the speech corpus that is used for training and evaluation. In this study, we propose BAlanced Script PROducer (BASPRO) system, which can automatically construct a phonetically balanced and rich set of Chinese sentences for collecting Mandarin Chinese speech data. First, we used pretrained natural language processing systems to extract ten-character candidate sentences from a large corpus of Chinese news texts. Then, we applied a genetic algorithm-based method to select 20 phonetically balanced sentence sets, each containing 20 sentences, from the candidate sentences. Using BASPRO, we obtained a recording script called TMNews, which contains 400 ten-character sentences. TMNews covers 84% of the syllables used in the real world. Moreover, the syllable distribution has 0.96 cosine similarity to the real-world syllable distribution. We converted the script into a speech corpus using two text-to-speech systems. Using the designed speech corpus, we tested the performances of speech enhancement (SE) and automatic speech recognition (ASR), which are one of the most important regression- and classification-based speech processing tasks, respectively. The experimental results show that the SE and ASR models trained on the designed speech corpus outperform their counterparts trained on a randomly composed speech corpus.

📄 PDF Abstract BibTeX arXiv:2301.04120

Code (1)

yuwchen/baspro 공식 구현 paddle

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)SentenceSpeech Enhancementspeech-recognitionSpeech Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Development and Transcription of Assamese Speech Corpus

2013-09-27 · Himangshu Sarma, Navanath Saharia, Utpal Sharma, Smriti Kumar Sinha 외

A balanced speech corpus is the basic need for any speech processing task. In this report we describe our effort on development of Assamese speech corpus. We mainly focused on some issues and challenges faced during deve…

Global Open Resources and Information for Language and Linguistic Analysis (GORILLA)

2016-05-01 · LREC 2016 5 · Damir Cavar, Malgorzata Cavar, Lwin Moe

The infrastructure Global Open Resources and Information for Language and Linguistic Analysis (GORILLA) was created as a resource that provides a bridge between disciplines such as documentary, theoretical, and corpus li…

AMISCO: The Austrian German Multi-Sensor Corpus

2016-05-01 · LREC 2016 5 · Hannes Pessentheiner, Thomas Pichler, Martin Hagm{\"u}ller

We introduce a unique, comprehensive Austrian German multi-sensor corpus with moving and non-moving speakers to facilitate the evaluation of estimators and detectors that jointly detect a speaker{'}s spatial and temporal…

GRASS: the Graz corpus of Read And Spontaneous Speech

2014-05-01 · LREC 2014 5 · Barbara Schuppler, Martin Hagmueller, Juan A. Morales-Cordovilla, Hannes Pessentheiner

This paper provides a description of the preparation, the speakers, the recordings, and the creation of the orthographic transcriptions of the first large scale speech database for Austrian German. It contains approximat…

Speech Recognition

A high quality and phonetic balanced speech corpus for Vietnamese

2019-04-11 · Pham Ngoc Phuong, Quoc Truong Do, Luong Chi Mai

This paper presents a high quality Vietnamese speech corpus that can be used for analyzing Vietnamese speech characteristic as well as building speech synthesis models. The corpus consists of 5400 clean-speech utterances…

Speech SynthesisVocal Bursts Intensity Prediction