paper-with-me

홈 › Papers

Enhancing Crowdsourced Audio for Text-to-Speech Models

2024-10-17 · José Giraldo, Martí Llopart-Font, Alex Peiró-Lilja, Carme Armentano-Oller, Gerard Sant, Baybars Külebi

High-quality audio data is a critical prerequisite for training robust text-to-speech models, which often limits the use of opportunistic or crowdsourced datasets. This paper presents an approach to overcome this limitation by implementing a denoising pipeline on the Catalan subset of Commonvoice, a crowd-sourced corpus known for its inherent noise and variability. The pipeline incorporates an audio enhancement phase followed by a selective filtering strategy. We developed an automatic filtering mechanism leveraging Non-Intrusive Speech Quality Assessment (NISQA) models to identify and retain the highest quality samples post-enhancement. To evaluate the efficacy of this approach, we trained a state of the art diffusion-based TTS model on the processed dataset. The results show a significant improvement, with an increase of 0.4 in the UTMOS Score compared to the baseline dataset without enhancement. This methodology shows promise for expanding the utility of crowdsourced data in TTS applications, particularly for mid to low resource languages like Catalan.

📄 PDF Abstract BibTeX arXiv:2410.13357

Code (0)

등록된 구현이 없습니다.

Tasks

Denoisingtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Crowdsourced Multilingual Speech Intelligibility Testing

2024-03-21 · Laura Lechler, Kamil Wojcicki

With the advent of generative audio features, there is an increasing need for rapid evaluation of their impact on speech intelligibility. Beyond the existing laboratory measures, which are expensive and do not scale well…

Speech Intelligibility Evaluation

Clotho: An Audio Captioning Dataset

2019-10-21 · Konstantinos Drossos, Samuel Lipping, Tuomas Virtanen

Audio captioning is the novel task of general audio content description using free text. It is an intermodal translation task (not speech-to-text), where a system accepts as an input an audio signal and outputs the textu…

Audio captioningDiversitySpeech-to-TextTranslation

CrowdSpeech and VoxDIY: Benchmark Datasets for Crowdsourced Audio Transcription

2021-07-02 · Nikita Pavlichenko, Ivan Stelmakh, Dmitry Ustalov

Domain-specific data is the crux of the successful transfer of machine learning systems from benchmarks to real life. In simple problems such as image classification, crowdsourcing has become one of the standard tools fo…

Crowdsourced Text Aggregationimage-classificationspeech-recognition

Evaluating the COVID-19 Identification ResNet (CIdeR) on the INTERSPEECH COVID-19 from Audio Challenges

2021-07-30 · Alican Akman, Harry Coppock, Alexander Gaskell, Panagiotis Tzirakis 외

We report on cross-running the recent COVID-19 Identification ResNet (CIdeR) on the two Interspeech 2021 COVID-19 diagnosis from cough and speech audio challenges: ComParE and DiCOVA. CIdeR is an end-to-end deep learning…

COVID-19 Diagnosis

Crowdsourcing and Evaluating Text-Based Audio Retrieval Relevances

2023-06-16 · Huang Xie, Khazar Khorrami, Okko Räsänen, Tuomas Virtanen

This paper explores grading text-based audio retrieval relevances with crowdsourcing assessments. Given a free-form text (e.g., a caption) as a query, crowdworkers are asked to grade audio clips using numeric scores (bet…

Audio captioningContrastive LearningRetrieval