paper-with-me

홈 › Papers

A Transfer Learning End-to-End ArabicText-To-Speech (TTS) Deep Architecture

2020-07-22 · Fady Fahmy, Mahmoud Khalil, Hazem Abbas

Speech synthesis is the artificial production of human speech. A typical text-to-speech system converts a language text into a waveform. There exist many English TTS systems that produce mature, natural, and human-like speech synthesizers. In contrast, other languages, including Arabic, have not been considered until recently. Existing Arabic speech synthesis solutions are slow, of low quality, and the naturalness of synthesized speech is inferior to the English synthesizers. They also lack essential speech key factors such as intonation, stress, and rhythm. Different works were proposed to solve those issues, including the use of concatenative methods such as unit selection or parametric methods. However, they required a lot of laborious work and domain expertise. Another reason for such poor performance of Arabic speech synthesizers is the lack of speech corpora, unlike English that has many publicly available corpora and audiobooks. This work describes how to generate high quality, natural, and human-like Arabic speech using an end-to-end neural deep network architecture. This work uses just $\langle$ text, audio $\rangle$ pairs with a relatively small amount of recorded audio samples with a total of 2.41 hours. It illustrates how to use English character embedding despite using diacritic Arabic characters as input and how to preprocess these audio samples to achieve the best results.

📄 PDF Abstract BibTeX arXiv:2007.11541

Code (0)

등록된 구현이 없습니다.

Tasks

RhythmSpeech Synthesistext-to-speechText to SpeechTransfer Learning

Similar Papers 제목 키워드 기반

ITAcotron 2: Transfering English Speech Synthesis Architectures and Speech Features to Italian

2021-11-01 · ICNLSP 2021 11 · Anna Favaro, Licia Sbattella, Roberto Tedesco, Vincenzo Scotti
Speech Synthesis

Effects of Layer Freezing on Transferring a Speech Recognition System to Under-resourced Languages

2021-02-08 · KONVENS (WS) 2021 9 · Onno Eberhard, Torsten Zesch

In this paper, we investigate the effect of layer freezing on the effectiveness of model transfer in the area of automatic speech recognition. We experiment with Mozilla's DeepSpeech architecture on German and Swiss Germ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Learned Transferable Architectures Can Surpass Hand-Designed Architectures for Large Scale Speech Recognition

2020-08-25 · Liqiang He, Dan Su, Dong Yu

In this paper, we explore the neural architecture search (NAS) for automatic speech recognition (ASR) systems. With reference to the previous works in the computer vision field, the transferability of the searched archit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)GPUNeural Architecture Search+2

Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation

2024-03-07 · Sai Akarsh, Vamshi Raghusimha, Anindita Mondal, Anil Vuppala

The language diversity in India's education sector poses a significant challenge, hindering inclusivity. Despite the democratization of knowledge through online educational content, the dominance of English, as the inter…

DiversityMachine Translationtext-to-speechText to Speech+1

A Transfer Learning Method for Speech Emotion Recognition from Automatic Speech Recognition

2020-08-06 · Sitong Zhou, Homayoon Beigi

This paper presents a transfer learning method in speech emotion recognition based on a Time-Delay Neural Network (TDNN) architecture. A major challenge in the current speech-based emotion detection research is data scar…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognition+3