paper-with-me

Papers

MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement

2025-11-15 · Xinyue Yu, Youqing Fang, Pingyu Wu, Guoyang Ye, Wenbo Zhou, Weiming Zhang, Song Xiao arxiv

Generating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speech factors and the coarse granularity of existing control mechanisms. To overcome these challenges, we have proposed a novel framework called MF-Speech, which consists of two core components: MF-SpeechEncoder and MF-SpeechGenerator. MF-SpeechEncoder acts as a factor purifier, adopting a multi-objective optimization strategy to decompose the original speech signal into highly pure and independent representations of content, timbre, and emotion. Subsequently, MF-SpeechGenerator functions as a conductor, achieving precise, composable and fine-grained control over these factors through dynamic fusion and Hierarchical Style Adaptive Normalization (HSAN). Experiments demonstrate that in the highly challenging multi-factor compositional speech generation task, MF-Speech significantly outperforms current state-of-the-art methods, achieving a lower word error rate (WER=4.67%), superior style control (SECS=0.5685, Corr=0.68), and the highest subjective evaluation scores(nMOS=3.96, sMOS_emotion=3.86, sMOS_style=3.78). Furthermore, the learned discrete factors exhibit strong transferability, demonstrating their significant potential as a general-purpose speech representation.

📄 PDF Abstract BibTeX arXiv:2511.12074

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data

2026-03-08 · Thanathai Lertpetchpun, Thanapat Trachu, Jihwan Lee, Tiantian Feng 외 arxiv

Accent is an integral part of society, reflecting multiculturalism and shaping how individuals express identity. The majority of English speakers are non-native (L2) speakers, yet current Text-To-Speech (TTS) systems pri…

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation

2026-07-09 · Shun Liu, Nan Xi, Yang Liu, Tianyu Luan 외 arxiv

Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understanding. Extending RIS to endoscopic imagery, however, presents unique c…

Image Segmentation

VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models

2025-04-04 · CVPR 2025 1 · Dahun Kim, AJ Piergiovanni, Ganesh Mallya, Anelia Angelova

We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing ben…

Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions

2024-03-25 · CVPR 2025 1 · Stefan Andreas Baumann, Felix Krause, Michael Neumayr, Nick Stracke 외

In recent years, advances in text-to-image (T2I) diffusion models have substantially elevated the quality of their generated images. However, achieving fine-grained control over attributes remains a challenge due to the …

Attribute

StylePTB: A Compositional Benchmark for Fine-grained Controllable Text Style Transfer

2021-04-12 · NAACL 2021 4 · Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy 외

Text style transfer aims to controllably generate text with targeted stylistic changes while maintaining core meaning from the source sentence constant. Many of the existing style transfer benchmarks primarily focus on i…

BenchmarkingSentenceStyle TransferText Generation+1