paper-with-me

Papers

Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion Model

2023-12-18 · Zhenyu Xie, Yang Wu, Xuehao Gao, Zhongqian Sun, Wei Yang, Xiaodan Liang

Text-guided motion synthesis aims to generate 3D human motion that not only precisely reflects the textual description but reveals the motion details as much as possible. Pioneering methods explore the diffusion model for text-to-motion synthesis and obtain significant superiority. However, these methods conduct diffusion processes either on the raw data distribution or the low-dimensional latent space, which typically suffer from the problem of modality inconsistency or detail-scarce. To tackle this problem, we propose a novel Basic-to-Advanced Hierarchical Diffusion Model, named B2A-HDM, to collaboratively exploit low-dimensional and high-dimensional diffusion models for high quality detailed motion synthesis. Specifically, the basic diffusion model in low-dimensional latent space provides the intermediate denoising result that to be consistent with the textual description, while the advanced diffusion model in high-dimensional latent space focuses on the following detail-enhancing denoising process. Besides, we introduce a multi-denoiser framework for the advanced diffusion model to ease the learning of high-dimensional model and fully explore the generative potential of the diffusion model. Quantitative and qualitative experiment results on two text-to-motion benchmarks (HumanML3D and KIT-ML) demonstrate that B2A-HDM can outperform existing state-of-the-art methods in terms of fidelity, modality consistency, and diversity.

📄 PDF Abstract BibTeX arXiv:2312.10960

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingMotion Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Making Social Platforms Accessible: Emotion-Aware Speech Generation with Integrated Text Analysis

2024-10-24 · Suparna De, Ionut Bostan, Nishanth Sastry

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to…

Speech Synthesistext-to-speechText to Speech

How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects

2025-03-06 · Wonkwang Lee, Jongwon Jeong, Taehong Moon, Hyeon-Jong Kim 외

Motion synthesis for diverse object categories holds great potential for 3D content creation but remains underexplored due to two key challenges: (1) the lack of comprehensive motion datasets that include a wide range of…

Motion Synthesis

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

2025-04-04 · CVPR 2025 1 · Ting-Hsuan Liao, Yi Zhou, Yu Shen, Chun-Hao Paul Huang 외

We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogen…

Language ModelingLanguage ModellingMotion GenerationMotion Synthesis+1

BiPO: Bidirectional Partial Occlusion Network for Text-to-Motion Synthesis

2024-11-28 · Seong-Eun Hong, Soobin Lim, Juyeong Hwang, Minwook Chang 외

Generating natural and expressive human motions from textual descriptions is challenging due to the complexity of coordinating full-body dynamics and capturing nuanced motion patterns over extended sequences that accurat…

Motion GenerationMotion Synthesis

SyntAct: A Synthesized Database of Basic Emotions

2022-06-01 · DCLRL (LREC) 2022 6 · Felix Burkhardt, Florian Eyben, Björn Schuller

Speech emotion recognition is in the focus of research since several decades and has many applications. One problem is sparse data for supervised learning. One way to tackle this problem is the synthesis of data with emo…

Emotion RecognitionSpeech Emotion RecognitionSpeech Synthesis