paper-with-me

Papers

T3M: Text Guided 3D Human Motion Synthesis from Speech

2024-08-23 · Wenshuo Peng, Kaipeng Zhang, Sai Qian Zhang

Speech-driven 3D motion synthesis seeks to create lifelike animations based on human speech, with potential uses in virtual reality, gaming, and the film production. Existing approaches reply solely on speech audio for motion generation, leading to inaccurate and inflexible synthesis results. To mitigate this problem, we introduce a novel text-guided 3D human motion synthesis method, termed \textit{T3M}. Unlike traditional approaches, T3M allows precise control over motion synthesis via textual input, enhancing the degree of diversity and user customization. The experiment results demonstrate that T3M can greatly outperform the state-of-the-art methods in both quantitative metrics and qualitative evaluations. We have publicly released our code at \href{https://github.com/Gloria2tt/T3M.git}{https://github.com/Gloria2tt/T3M.git}

📄 PDF Abstract BibTeX arXiv:2408.12885

Code (1)

gloria2tt/t3m 공식 구현 pytorch

Tasks

DiversityMotion GenerationMotion Synthesis

Similar Papers 제목 키워드 기반

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

2025-06-03 · Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse…

Expressive Speech SynthesisPrompt LearningSpeech Synthesistext-to-speech+1

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

2026-07-17 · Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu 외 arxiv

Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions…

Speech Synthesis

DeepGesture: A conversational gesture synthesis system based on emotions and semantics

2025-07-03 · Thanh Hoang-Minh

Along with the explosion of large language models, improvements in speech synthesis, advancements in hardware, and the evolution of computer graphics, the current bottleneck in creating digital humans lies in generating …

Gesture GenerationMotion SynthesisSpeech SynthesisUnity

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

2026-06-08 · Minghui Wu, Ganjun Liu, Zikun Fang, Ting Meng 외 arxiv

Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained inte…

Speech Synthesis

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control

2025-01-10 · Shaozuo Zhang, Ambuj Mehrish, Yingting Li, Soujanya Poria

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuanc…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis