paper-with-me

Papers

Generating Holistic 3D Human Motion from Speech

2022-12-08 · CVPR 2023 1 · Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, DaCheng Tao, Michael J. Black

This work addresses the problem of generating 3D holistic body motions from human speech. Given a speech recording, we synthesize sequences of 3D body poses, hand gestures, and facial expressions that are realistic and diverse. To achieve this, we first build a high-quality dataset of 3D holistic body meshes with synchronous speech. We then define a novel speech-to-motion generation framework in which the face, body, and hands are modeled separately. The separated modeling stems from the fact that face articulation strongly correlates with human speech, while body poses and hand gestures are less correlated. Specifically, we employ an autoencoder for face motions, and a compositional vector-quantized variational autoencoder (VQ-VAE) for the body and hand motions. The compositional VQ-VAE is key to generating diverse results. Additionally, we propose a cross-conditional autoregressive model that generates body poses and hand gestures, leading to coherent and realistic motions. Extensive experiments and user studies demonstrate that our proposed approach achieves state-of-the-art performance both qualitatively and quantitatively. Our novel dataset and code will be released for research purposes at https://talkshow.is.tue.mpg.de.

📄 PDF Abstract BibTeX arXiv:2212.04420

Code (3)

yhw-yhw/talkshow 공식 구현 pytorch
yhw-yhw/show pytorch
zhenglinzhou/headstudio jax

Tasks

3D Face AnimationGesture GenerationMotion Generation

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…

Similar Papers 제목 키워드 기반

3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control

2026-01-26 · Xuanmeng Sha, Liyun Zhang, Tomohiro Mashita, Naoya Chiba 외 arxiv

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing…

Gesture Generation

Combo: Co-speech holistic 3D human motion generation and efficient customizable adaptation in harmony

2024-08-18 · Chao Xu, Mingze Sun, Zhi-Qi Cheng, Fei Wang 외

In this paper, we propose a novel framework, Combo, for harmonious co-speech holistic 3D human motion generation and efficient customizable adaption. In particular, we identify that one fundamental challenge as the multi…

Motion Generationparameter-efficient fine-tuning

Towards Variable and Coordinated Holistic Co-Speech Motion Generation

2024-03-30 · CVPR 2024 1 · Yifei Liu, Qiong Cao, Yandong Wen, Huaiguang Jiang 외

This paper addresses the problem of generating lifelike holistic co-speech motions for 3D avatars, focusing on two key aspects: variability and coordination. Variability allows the avatar to exhibit a wide range of motio…

Motion GenerationQuantization

Holistic-Motion2D: Scalable Whole-body Human Motion Generation in 2D Space

2024-06-17 · YuAn Wang, Zhao Wang, Junhao Gong, Di Huang 외

In this paper, we introduce a novel path to $\textit{general}$ human motion generation by focusing on 2D space. Traditional methods have primarily generated human motions in 3D, which, while detailed and realistic, are o…

Motion Generation

EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling

2023-12-31 · CVPR 2024 1 · Haiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng 외

We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements. To achieve this, we first introduce BEAT2 (BEAT-SMPLX-FLAME…

3D Face AnimationDiversityGesture GenerationRhythm