paper-with-me

Papers

AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data

2023-09-25 · Jianwei Yu, Hangting Chen, Yanyao Bian, Xiang Li, Yi Luo, Jinchuan Tian, Mengyang Liu, Jiayi Jiang, Shuai Wang

Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained. One reason for this limitation is that a considerable amount of the publicly available speech data is compromised by background noise, speech overlapping, lack of speech segmentation information, missing speaker labels, and incomplete transcriptions, which can largely hinder their usefulness. On the other hand, human annotation of speech data is both time-consuming and costly. To address this issue, we introduce an automatic in-the-wild speech data preprocessing framework (AutoPrep) in this paper, which is designed to enhance speech quality, generate speaker labels, and produce transcriptions automatically. The proposed AutoPrep framework comprises six components: speech enhancement, speech segmentation, speaker clustering, target speech extraction, quality filtering and automatic speech recognition. Experiments conducted on the open-sourced WenetSpeech and our self-collected AutoPrepWild corpora demonstrate that the proposed AutoPrep framework can generate preprocessed data with similar DNSMOS and PDNSMOS scores compared to several open-sourced TTS datasets. The corresponding TTS system can achieve up to 0.68 in-domain speaker similarity.

📄 PDF Abstract BibTeX arXiv:2309.13905

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionSpeech EnhancementSpeech Extractionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework

2024-12-10 · Meihao Fan, Ju Fan, Nan Tang, Lei Cao 외

Answering natural language (NL) questions about tables, known as Tabular Question Answering (TQA), is crucial because it allows users to quickly and efficiently extract meaningful insights from structured data, effective…

Code GenerationLarge Language ModelQuestion Answering

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations

2025-05-08 · Linrong Pan, Chenglong Jiang, Gaoze Hou, Ying Gao

This paper reports the construction of the Teochew-Wild, a speech corpus of the Teochew dialect. The corpus includes 18.9 hours of in-the-wild Teochew speech data from multiple speakers, covering both formal and colloqui…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

KoSpeech: Open-Source Toolkit for End-to-End Korean Speech Recognition

2020-09-07 · Soohwan Kim, Seyoung Bae, Cheolhwang Won

We present KoSpeech, an open-source software, which is modular and extensible end-to-end Korean automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch. Several automatic speech recognition …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

2025-01-27 · Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li 외

Recent advancements in speech generation have been driven by the large-scale training datasets. However, current models fall short of capturing the spontaneity and variability inherent in real-world human speech, due to …

WildSpoof Challenge Evaluation Plan

2025-08-23 · Yihan Wu, Jee-weon Jung, Hye-jin Shim, Xin Cheng 외 arxiv

The WildSpoof Challenge aims to advance the use of in-the-wild data in two intertwined speech processing tasks. It consists of two parallel tracks: (1) Text-to-Speech (TTS) synthesis for generating spoofed speech, and (2…

Speaker Verification