paper-with-me

Papers

Phonological Features for 0-shot Multilingual Speech Synthesis

2020-08-06 · Marlene Staib, Tian Huey Teh, Alexandra Torresquintero, Devang S Ram Mohan, Lorenzo Foglianti, Raphael Lenain, Jiameng Gao

Code-switching---the intra-utterance use of multiple languages---is prevalent across the world. Within text-to-speech (TTS), multilingual models have been found to enable code-switching. By modifying the linguistic input to sequence-to-sequence TTS, we show that code-switching is possible for languages unseen during training, even within monolingual models. We use a small set of phonological features derived from the International Phonetic Alphabet (IPA), such as vowel height and frontness, consonant place and manner. This allows the model topology to stay unchanged for different languages, and enables new, previously unseen feature combinations to be interpreted by the model. We show that this allows us to generate intelligible, code-switched speech in a new language at test time, including the approximation of sounds never seen in training.

📄 PDF Abstract BibTeX arXiv:2008.04107

Code (1)

papercup-open-source/phonological-features 공식 구현

Tasks

Speech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Multilingual and crosslingual speech recognition using phonological-vector based phone embeddings

2021-07-11 · Chengrui Zhu, Keyu An, Huahuan Zheng, Zhijian Ou

The use of phonological features (PFs) potentially allows language-specific phones to remain linked in training, which is highly desirable for information sharing for multilingual and crosslingual speech recognition meth…

speech-recognitionSpeech Recognition

Cross-lingual Low Resource Speaker Adaptation Using Phonological Features

2021-11-17 · Georgia Maniati, Nikolaos Ellinas, Konstantinos Markopoulos, Georgios Vamvoukakis 외

The idea of using phonological features instead of phonemes as input to sequence-to-sequence TTS has been recently proposed for zero-shot multilingual speech synthesis. This approach is useful for code-switching, as it f…

Speech Synthesis

Applying Feature Underspecified Lexicon Phonological Features in Multilingual Text-to-Speech

2022-04-14 · Cong Zhang, Huinan Zeng, Huang Liu, Jiewen Zheng

This study investigates whether the phonological features derived from the Featurally Underspecified Lexicon model can be applied in text-to-speech systems to generate native and non-native speech in English and Mandarin…

Language Acquisitiontext-to-speechText to Speech

Applying Phonological Features in Multilingual Text-To-Speech

2021-10-07 · Cong Zhang, Huinan Zeng, Huang Liu, Jiewen Zheng

This study investigates whether phonological features can be applied in text-to-speech systems to generate native and non-native speech in English and Mandarin. We present a mapping of ARPABET/pinyin to SAMPA/SAMPA-SC an…

Language Acquisitiontext-to-speechText to Speech

Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

2026-05-25 · Abner Hernandez, Tomás Arias-Vergara, Daiqi Liu, Andreas Maier 외 arxiv

Phonological features provide a language-general and linguistically grounded representation of speech. We present PhonoQ-2.0, a multilingual frame-level phonological feature recognizer built on self-supervised speech mod…