paper-with-me

Papers

ReverBERT: A State Space Model for Efficient Text-Driven Speech Style Transfer

2025-03-26 · Michael Brown, Sofia Martinez, Priya Singh

Text-driven speech style transfer aims to mold the intonation, pace, and timbre of a spoken utterance to match stylistic cues from text descriptions. While existing methods leverage large-scale neural architectures or pre-trained language models, the computational costs often remain high. In this paper, we present \emph{ReverBERT}, an efficient framework for text-driven speech style transfer that draws inspiration from a state space model (SSM) paradigm, loosely motivated by the image-based method of Wang and Liu~\cite{wang2024stylemamba}. Unlike image domain techniques, our method operates in the speech space and integrates a discrete Fourier transform of latent speech features to enable smooth and continuous style modulation. We also propose a novel \emph{Transformer-based SSM} layer for bridging textual style descriptors with acoustic attributes, dramatically reducing inference time while preserving high-quality speech characteristics. Extensive experiments on benchmark speech corpora demonstrate that \emph{ReverBERT} significantly outperforms baselines in terms of naturalness, expressiveness, and computational efficiency. We release our model and code publicly to foster further research in text-driven speech style transfer.

📄 PDF Abstract BibTeX arXiv:2503.20992

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyStyle Transfer

Similar Papers 제목 키워드 기반

Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation

2025-07-25 · Fang Kang, Yin Cao, Haoyu Chen arxiv

Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more cha…

Talking Face GenerationFace Alignment

Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation

2023-12-18 · Hui Fu, Zeqing Wang, Ke Gong, Keze Wang 외

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip s…

DisentanglementRepresentation Learning

UniFLG: Unified Facial Landmark Generator from Text or Speech

2023-02-28 · Kentaro Mitsui, Yukiya Hono, Kei Sawada

Talking face generation has been extensively investigated owing to its wide applicability. The two primary frameworks used for talking face generation comprise a text-driven framework, which generates synchronized speech…

DecoderFace GenerationSpeech SynthesisTalking Face Generation+2

Generating coherent spontaneous speech and gesture from text

2021-01-14 · Simon Alexanderson, Éva Székely, Gustav Eje Henter, Taras Kucherenko 외

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating…

Gesture GenerationMotion GenerationSpeech Synthesistext-to-speech+1

Neural Text to Articulate Talk: Deep Text to Audiovisual Speech Synthesis achieving both Auditory and Photo-realism

2023-12-11 · Georgios Milis, Panagiotis P. Filntisis, Anastasios Roussos, Petros Maragos

Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-sync…

Face GenerationLip ReadingSpeech SynthesisTalking Face Generation+2