paper-with-me

홈 › Papers

MCM: Multi-condition Motion Synthesis Framework

2024-04-19 · Zeyu Ling, Bo Han, Yongkang Wongkan, Han Lin, Mohan Kankanhalli, Weidong Geng

Conditional human motion synthesis (HMS) aims to generate human motion sequences that conform to specific conditions. Text and audio represent the two predominant modalities employed as HMS control conditions. While existing research has primarily focused on single conditions, the multi-condition human motion synthesis remains underexplored. In this study, we propose a multi-condition HMS framework, termed MCM, based on a dual-branch structure composed of a main branch and a control branch. This framework effectively extends the applicability of the diffusion model, which is initially predicated solely on textual conditions, to auditory conditions. This extension encompasses both music-to-dance and co-speech HMS while preserving the intrinsic quality of motion and the capabilities for semantic association inherent in the original model. Furthermore, we propose the implementation of a Transformer-based diffusion model, designated as MWNet, as the main branch. This model adeptly apprehends the spatial intricacies and inter-joint correlations inherent in motion sequences, facilitated by the integration of multi-wise self-attention modules. Extensive experiments show that our method achieves competitive results in single-condition and multi-condition HMS tasks.

📄 PDF Abstract BibTeX arXiv:2404.12886

Code (1)

fluide1022/MCM pytorch

Tasks

Motion Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis

2024-05-24 · Ke Fan, Junshu Tang, Weijian Cao, Ran Yi 외

Text-to-motion synthesis is a crucial task in computer vision. Existing methods are limited in their universality, as they are tailored for single-person or two-person scenarios and can not be applied to generate motions…

Motion GenerationMotion Synthesis

Dynamic Motion Synthesis: Masked Audio-Text Conditioned Spatio-Temporal Transformers

2024-09-03 · Sohan Anisetty, James Hays

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized…

Language ModelingLanguage ModellingMasked Language ModelingMotion Generation+1

MCM: Multi-condition Motion Synthesis Framework for Multi-scenario

2023-09-06 · Zeyu Ling, Bo Han, Yongkang Wong, Mohan Kangkanhalli 외

The objective of the multi-condition human motion synthesis task is to incorporate diverse conditional inputs, encompassing various forms like text, music, speech, and more. This endows the task with the capability to ad…

Motion GenerationMotion Synthesis

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

2026-05-28 · Yiheng Li, Zhuo Li, Ruibing Hou, Yingjie Chen 외 arxiv

Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific …

Motion Synthesis

GUESS:GradUally Enriching SyntheSis for Text-Driven Human Motion Generation

2024-01-04 · Xuehao Gao, Yang Yang, Zhenyu Xie, Shaoyi Du 외

In this paper, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strat…

Motion GenerationMotion Synthesis