paper-with-me

홈 › Papers

Intelli-Z: Toward Intelligible Zero-Shot TTS

2024-01-25 · Sunghee Jung, Won Jang, Jaesam Yoon, BongWan Kim

Although numerous recent studies have suggested new frameworks for zero-shot TTS using large-scale, real-world data, studies that focus on the intelligibility of zero-shot TTS are relatively scarce. Zero-shot TTS demands additional efforts to ensure clear pronunciation and speech quality due to its inherent requirement of replacing a core parameter (speaker embedding or acoustic prompt) with a new one at the inference stage. In this study, we propose a zero-shot TTS model focused on intelligibility, which we refer to as Intelli-Z. Intelli-Z learns speaker embeddings by using multi-speaker TTS as its teacher and is trained with a cycle-consistency loss to include mismatched text-speech pairs for training. Additionally, it selectively aggregates speaker embeddings along the temporal dimension to minimize the interference of the text content of reference speech at the inference stage. We substantiate the effectiveness of the proposed methods with an ablation study. The Mean Opinion Score (MOS) increases by 9% for unseen speakers when the first two methods are ap- plied, and it further improves by 16% when selective temporal aggregation is applied.

📄 PDF Abstract BibTeX arXiv:2401.13921

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Zero-shot Generative Linguistic Steganography

2024-03-16 · Ke Lin, Yiyang Luo, Zijian Zhang, Ping Luo

Generative linguistic steganography attempts to hide secret messages into covertext. Previous studies have generally focused on the statistical differences between the covertext and stegotext, however, ill-formed stegote…

In-Context LearningLinguistic steganography

Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining

2023-01-30 · Takaaki Saeki, Soumi Maiti, Xinjian Li, Shinji Watanabe 외

While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This pape…

Language ModelingLanguage Modellingtext-to-speechText to Speech

An Evaluation Framework for Text-to-Speech Voice Reconstruction

2026-06-19 · Ariadna Sanchez, Christoph Minixhofer, Korin Richmond, Ondrej Klejch 외 arxiv

Voice reconstruction using Text-to-Speech (TTS) offers a communication method for people with speech disorders, which aims to retain their speaker identity while improving intelligibility. Previous work generally relies …

Coding Speech through Vocal Tract Kinematics

2024-06-18 · Cheol Jun Cho, Peter Wu, Tejas S. Prabhune, Dhruv Agarwal 외

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spo…

Voice Conversion

Are Mutually Intelligible Languages Easier to Translate?

2022-01-31 · Avital Friedland, Jonathan Zeltser, Omer Levy

Two languages are considered mutually intelligible if their native speakers can communicate with each other, while using their own mother tongue. How does the fact that humans perceive a language pair as mutually intelli…

Translation