paper-with-me

홈 › Papers

Controllable Emphasis with zero data for text-to-speech

2023-07-13 · Arnaud Joly, Marco Nicolis, Ekaterina Peterova, Alessandro Lombardi, Ammar Abbas, Arent van Korlaar, Aman Hussain, Parul Sharma, Alexis Moinet, Mateusz Lajszczak, Penny Karanasou, Antonio Bonafonte, Thomas Drugman, Elena Sokolova

We present a scalable method to produce high quality emphasis for text-to-speech (TTS) that does not require recordings or annotations. Many TTS models include a phoneme duration model. A simple but effective method to achieve emphasized speech consists in increasing the predicted duration of the emphasised word. We show that this is significantly better than spectrogram modification techniques improving naturalness by $7.3\%$ and correct testers' identification of the emphasized word in a sentence by $40\%$ on a reference female en-US voice. We show that this technique significantly closes the gap to methods that require explicit recordings. The method proved to be scalable and preferred in all four languages tested (English, Spanish, Italian, German), for different voices and multiple speaking styles.

📄 PDF Abstract BibTeX arXiv:2307.07062

Code (0)

등록된 구현이 없습니다.

Tasks

Sentencetext-to-speechText to Speech

Similar Papers 제목 키워드 기반

StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation

2025-10-15 · Xi Chen, Yuchen Song, Satoshi Nakamura arxiv

We propose a stress-aware speech-to-speech translation (S2ST) system that preserves word-level emphasis by leveraging LLMs for cross-lingual emphasis conversion. Our method translates source-language stress into target-l…

Speech-to-Speech Translation

ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models

2023-05-23 · Minki Kang, Wooseok Han, Sung Ju Hwang, Eunho Yang

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Zero-Bit Transmission of Adaptive Pre- and De-emphasis Filters for Speech and Audio Coding

2024-07-02 · Niloofar Omidi Piralideh, Philippe Gournay, Roch Lefebvre

This paper introduces a novel adaptation approach for first-order pre- and de-emphasis filters, an essential tool in many speech and audio codecs to increase coding efficiency and perceived quality. The proposed zero-bit…

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

2025-02-08 · Wei Deng, Siyi Zhou, Jingchen Shu, Jinchao Wang 외

Recently, large language model (LLM) based text-to-speech (TTS) systems have gradually become the mainstream in the industry due to their high naturalness and powerful zero-shot voice cloning capabilities.Here, we introd…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+4

ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

2024-06-03 · Shengpeng Ji, Qian Chen, Wen Wang, Jialong Zuo 외

In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker's voice and enabling arbitrary control and adjustment of speaking style. Prior zero-shot TTS models only mimic t…

Speech Synthesistext-to-speechText to Speech