paper-with-me

홈 › Papers

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models

2026-04-09 · Jaehoon Kang, Yejin Lee, Yoonji Park, Kyuhong Shim arxiv

While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control and apply a single global style across an utterance. This restricts practical use cases that require continuous style attribute interpolation across utterances and time-varying style transitions within a single utterance. In this paper, we propose novel techniques to achieve both capabilities in existing prompt-based TTS models. For inter-utterance style interpolation, we compute direction vectors between contrastive style prompts in the embedding space and perform simple interpolation, enabling smooth transitions between style characteristics. For intra-utterance style transition, we first identify a strong attention bias toward early tokens in autoregressive TTS decoders, causing the initial audio realization to dominate subsequent generation. To mitigate this effect, we introduce KV-cache swapping and sliding-window attention masking. Experiments demonstrate that our proposed inter-utterance interpolation achieves a 99-100% success rate in gender conversion, up to 36 Hz pitch variation, and up to 1.6 syllables-per-second speed change. Our intra-utterance transition maintains a speaker similarity of 0.81-0.91 and achieves perceptual smoothness scores of 3.48-4.48.

📄 PDF Abstract BibTeX arXiv:2605.27376

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hierarchical Multi-Grained Generative Model for Expressive Speech Synthesis

2020-09-17 · Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto 외

This paper proposes a hierarchical generative model with a multi-grained latent variable to synthesize expressive speech. In recent years, fine-grained latent variables are introduced into the text-to-speech synthesis th…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech+1

Learning utterance-level representations through token-level acoustic latents prediction for Expressive Speech Synthesis

2022-11-01 · Karolos Nikitaras, Konstantinos Klapsas, Nikolaos Ellinas, Georgia Maniati 외

This paper proposes an Expressive Speech Synthesis model that utilizes token-level latent prosodic variables in order to capture and control utterance-level attributes, such as character acting voice and speaking style. …

DisentanglementDiversityExpressive Speech SynthesisSpeech Synthesis

Customized Conversational Recommender Systems

2022-06-30 · Shuokai Li, Yongchun Zhu, Ruobing Xie, Zhenwei Tang 외

Conversational recommender systems (CRS) aim to capture user's current intentions and provide recommendations through real-time multi-turn conversational interactions. As a human-machine interactive system, it is essenti…

Meta-LearningRecommendation Systems

Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis

2021-09-08 · Songxiang Liu, Shan Yang, Dan Su, Dong Yu

Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker's voice. Most previous CSST approaches rely on expensive high-qual…

Expressive Speech SynthesisSentenceSpeech SynthesisStyle Transfer+2

Hierarchical Generative Modeling for Controllable Speech Synthesis

2018-10-16 · ICLR 2019 5 · Wei-Ning Hsu, Yu Zhang, Ron J. Weiss, Heiga Zen 외

This paper proposes a neural sequence-to-sequence text-to-speech (TTS) model which can control latent attributes in the generated speech that are rarely annotated in the training data, such as speaking style, accent, bac…

AttributeSpeech Synthesistext-to-speechText to Speech