paper-with-me

Papers

Unsupervised Data Selection for TTS: Using Arabic Broadcast News as a Case Study

2023-01-22 · Massa Baali, Tomoki Hayashi, Hamdy Mubarak, Soumi Maiti, Shinji Watanabe, Wassim El-Hajj, Ahmed Ali

Several high-resource Text to Speech (TTS) systems currently produce natural, well-established human-like speech. In contrast, low-resource languages, including Arabic, have very limited TTS systems due to the lack of resources. We propose a fully unsupervised method for building TTS, including automatic data selection and pre-training/fine-tuning strategies for TTS training, using broadcast news as a case study. We show how careful selection of data, yet smaller amounts, can improve the efficiency of TTS system in generating more natural speech than a system trained on a bigger dataset. We adopt to propose different approaches for the: 1) data: we applied automatic annotations using DNSMOS, automatic vowelization, and automatic speech recognition (ASR) for fixing transcriptions' errors; 2) model: we used transfer learning from high-resource language in TTS model and fine-tuned it with one hour broadcast recording then we used this model to guide a FastSpeech2-based Conformer model for duration. Our objective evaluation shows 3.9% character error rate (CER), while the groundtruth has 1.3% CER. As for the subjective evaluation, where 1 is bad and 5 is excellent, our FastSpeech2-based Conformer model achieved a mean opinion score (MOS) of 4.4 for intelligibility and 4.2 for naturalness, where many annotators recognized the voice of the broadcaster, which proves the effectiveness of our proposed unsupervised method.

📄 PDF Abstract BibTeX arXiv:2301.09099

Code (2)

espnet/espnet/tree/master/egs2/qasr_tts/tts1 공식 구현 pytorch
MindCode-4/code-3/tree/main/fastspeech2_conformer mindspore

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognitiontext-to-speechText to SpeechTransfer Learning

Similar Papers 제목 키워드 기반

Expanding Arabic Treebank to Speech: Results from Broadcast News

2012-05-01 · LREC 2012 5 · Mohamed Maamouri, Ann Bies, Seth Kulick

Treebanking a large corpus of relatively structured speech transcribed from various Arabic Broadcast News (BN) sources has allowed us to begin to address the many challenges of annotating and parsing a speech corpus in A…

Morphological Analysis

End-to-End Speech Translation of Arabic to English Broadcast News

2022-12-11 · Fethi Bougares, Salim Jouili

Speech translation (ST) is the task of directly translating acoustic speech signals in a source language into text in a foreign language. ST task has been addressed, for a long time, using a pipeline approach with two mo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMachine Translation+4

Unsupervised Broadcast News Summarization; a comparative study on Maximal Marginal Relevance (MMR) and Latent Semantic Analysis (LSA)

2023-01-05 · Majid Ramezani, Mohammad-Salar Shahryari, Amir-Reza Feizi-Derakhshi, Mohammad-Reza Feizi-Derakhshi

The methods of automatic speech summarization are classified into two groups: supervised and unsupervised methods. Supervised methods are based on a set of features, while unsupervised methods perform summarization based…

News Summarization

OSIAN: Open Source International Arabic News Corpus - Preparation and Integration into the CLARIN-infrastructure

2019-08-01 · WS 2019 8 · Imad Zeroual, Dirk Goldhahn, Thomas Eckart, Abdelhak Lakhouaja

The World Wide Web has become a fundamental resource for building large text corpora. Broadcasting platforms such as news websites are rich sources of data regarding diverse topics and form a valuable foundation for rese…

ArticlesDescriptiveLEMMA

Linear Semantic Segmentation for Low-Resource Spoken Dialects

2026-05-07 · Kirill Chirkunov, Younes Samih, Abed Alhakim Freihat, Hanan Aldarmaki arxiv

Semantic segmentation is a core component of discourse analysis, yet existing models are primarily developed and evaluated on high-resource written text, limiting their effectiveness on low-resource spoken varieties. In …

Semantic Segmentation