paper-with-me

홈 › Papers

Fine-grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement

2020-11-08 · Daxin Tan, Tan Lee

This paper presents a novel design of neural network system for fine-grained style modeling, transfer and prediction in expressive text-to-speech (TTS) synthesis. Fine-grained modeling is realized by extracting style embeddings from the mel-spectrograms of phone-level speech segments. Collaborative learning and adversarial learning strategies are applied in order to achieve effective disentanglement of content and style factors in speech and alleviate the "content leakage" problem in style modeling. The proposed system can be used for varying-content speech style transfer in the single-speaker scenario. The results of objective and subjective evaluation show that our system performs better than other fine-grained speech style transfer models, especially in the aspect of content preservation. By incorporating a style predictor, the proposed system can also be used for text-to-speech synthesis. Audio samples are provided for system demonstration https://daxintan-cuhk.github.io/pl-csd-speech .

📄 PDF Abstract BibTeX arXiv:2011.03943

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementSpeech SynthesisStyle Transfertext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

StylePTB: A Compositional Benchmark for Fine-grained Controllable Text Style Transfer

2021-04-12 · NAACL 2021 4 · Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy 외

Text style transfer aims to controllably generate text with targeted stylistic changes while maintaining core meaning from the source sentence constant. Many of the existing style transfer benchmarks primarily focus on i…

BenchmarkingSentenceStyle TransferText Generation+1

Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis

2023-03-14 · Chunyu Qiang, Peng Yang, Hao Che, Ying Zhang 외

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker's timbre. In most previous methods, the synthesized fine-grained prosody feature…

Prosody PredictionSpeech SynthesisStyle Transfer

Fine-Grained Image Style Transfer with Visual Transformers

2022-10-11 · Jianbo Wang, Huan Yang, Jianlong Fu, Toshihiko Yamasaki 외

With the development of the convolutional neural network, image style transfer has drawn increasing attention. However, most existing approaches adopt a global feature transformation to transfer style patterns into conte…

Style Transfer

One-shot Embroidery Customization via Contrastive LoRA Modulation

2025-09-23 · Jun Ma, Qian He, Gaofeng He, Huang Chen 외 arxiv

Diffusion models have significantly advanced image manipulation techniques, and their ability to generate photorealistic images is beginning to transform retail workflows, particularly in presale visualization. Beyond ar…

Knowledge DistillationContrastive LearningImage ManipulationStyle Transfer

Fine-Grained Emotion Prediction by Modeling Emotion Definitions

2021-07-26 · Gargi Singh, Dhanajit Brahma, Piyush Rai, Ashutosh Modi

In this paper, we propose a new framework for fine-grained emotion prediction in the text through emotion definition modeling. Our approach involves a multi-task learning framework that models definitions of emotions as …

Language ModelingLanguage ModellingMasked Language ModelingMulti-Task Learning+2