paper-with-me

홈 › Papers

Strategies in Transfer Learning for Low-Resource Speech Synthesis: Phone Mapping, Features Input, and Source Language Selection

2023-06-21 · Phat Do, Matt Coler, Jelske Dijkstra, Esther Klabbers

We compare using a PHOIBLE-based phone mapping method and using phonological features input in transfer learning for TTS in low-resource languages. We use diverse source languages (English, Finnish, Hindi, Japanese, and Russian) and target languages (Bulgarian, Georgian, Kazakh, Swahili, Urdu, and Uzbek) to test the language-independence of the methods and enhance the findings' applicability. We use Character Error Rates from automatic speech recognition and predicted Mean Opinion Scores for evaluation. Results show that both phone mapping and features input improve the output quality and the latter performs better, but these effects also depend on the specific language combination. We also compare the recently-proposed Angular Similarity of Phone Frequencies (ASPF) with a family tree-based distance measure as a criterion to select source languages in transfer learning. ASPF proves effective if label-based phone input is used, while the language distance does not have expected effects.

📄 PDF Abstract BibTeX arXiv:2306.12040

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech SynthesisTransfer Learning

Similar Papers 제목 키워드 기반

Fine-grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement

2020-11-08 · Daxin Tan, Tan Lee

This paper presents a novel design of neural network system for fine-grained style modeling, transfer and prediction in expressive text-to-speech (TTS) synthesis. Fine-grained modeling is realized by extracting style emb…

DisentanglementSpeech SynthesisStyle Transfertext-to-speech+2

Speech Synthesis for Low Resource Languages using Transliteration Enabled Transfer Learning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In the area of Human Computer Interaction (HCI), Text To Speech (TTS) synthesis has received a significant boost in recent years, especially with the development of various deep learning techniques capable of generating …

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+3

Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios

2025-05-30 · Gerard I. Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori 외

We propose a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing ph…

Cross-Lingual TransferPhoneme RecognitionSpeech-to-TextSpeech-to-Text Translation+1

Learning pronunciation from a foreign language in speech synthesis networks

2018-11-23 · Young-Gun Lee, Suwon Shon, Taesu Kim

Although there are more than 6,500 languages in the world, the pronunciations of many phonemes sound similar across the languages. When people learn a foreign language, their pronunciation often reflects their native lan…

Speech Synthesis

Towards Multi-Scale Style Control for Expressive Speech Synthesis

2021-04-08 · Xiang Li, Changhe Song, Jingbei Li, Zhiyong Wu 외

This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis. The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level an…

Expressive Speech SynthesisSpeech SynthesisStyle Transfer