paper-with-me

홈 › Papers

MCM: Multi-condition Motion Synthesis Framework for Multi-scenario

2023-09-06 · Zeyu Ling, Bo Han, Yongkang Wong, Mohan Kangkanhalli, Weidong Geng

The objective of the multi-condition human motion synthesis task is to incorporate diverse conditional inputs, encompassing various forms like text, music, speech, and more. This endows the task with the capability to adapt across multiple scenarios, ranging from text-to-motion and music-to-dance, among others. While existing research has primarily focused on single conditions, the multi-condition human motion generation remains underexplored. In this paper, we address these challenges by introducing MCM, a novel paradigm for motion synthesis that spans multiple scenarios under diverse conditions. The MCM framework is able to integrate with any DDPM-like diffusion model to accommodate multi-conditional information input while preserving its generative capabilities. Specifically, MCM employs two-branch architecture consisting of a main branch and a control branch. The control branch shares the same structure as the main branch and is initialized with the parameters of the main branch, effectively maintaining the generation ability of the main branch and supporting multi-condition input. We also introduce a Transformer-based diffusion model MWNet (DDPM-like) as our main branch that can capture the spatial complexity and inter-joint correlations in motion sequences through a channel-dimension self-attention module. Quantitative comparisons demonstrate that our approach achieves SoTA results in both text-to-motion and competitive results in music-to-dance tasks, comparable to task-specific methods. Furthermore, the qualitative evaluation shows that MCM not only streamlines the adaptation of methodologies originally designed for text-to-motion tasks to domains like music-to-dance and speech-to-gesture, eliminating the need for extensive network re-configurations but also enables effective multi-condition modal control, realizing "once trained is motion need".

📄 PDF Abstract BibTeX arXiv:2309.03031

Code (1)

fluide1022/MCM pytorch

Tasks

Motion GenerationMotion Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis

2024-05-24 · Ke Fan, Junshu Tang, Weijian Cao, Ran Yi 외

Text-to-motion synthesis is a crucial task in computer vision. Existing methods are limited in their universality, as they are tailored for single-person or two-person scenarios and can not be applied to generate motions…

Motion GenerationMotion Synthesis

MCM: Multi-condition Motion Synthesis Framework

2024-04-19 · Zeyu Ling, Bo Han, Yongkang Wongkan, Han Lin 외

Conditional human motion synthesis (HMS) aims to generate human motion sequences that conform to specific conditions. Text and audio represent the two predominant modalities employed as HMS control conditions. While exis…

Motion Synthesis

Dynamic Motion Synthesis: Masked Audio-Text Conditioned Spatio-Temporal Transformers

2024-09-03 · Sohan Anisetty, James Hays

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized…

Language ModelingLanguage ModellingMasked Language ModelingMotion Generation+1

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

2026-05-28 · Yiheng Li, Zhuo Li, Ruibing Hou, Yingjie Chen 외 arxiv

Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific …

Motion Synthesis

GUESS:GradUally Enriching SyntheSis for Text-Driven Human Motion Generation

2024-01-04 · Xuehao Gao, Yang Yang, Zhenyu Xie, Shaoyi Du 외

In this paper, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strat…

Motion GenerationMotion Synthesis