paper-with-me

홈 › Papers

Fleximo: Towards Flexible Text-to-Human Motion Video Generation

2024-11-29 · Yuhang Zhang, Yuan Zhou, Zeyu Liu, Yuxuan Cai, Qiuyue Wang, Aidong Men, Huan Yang

Current methods for generating human motion videos rely on extracting pose sequences from reference videos, which restricts flexibility and control. Additionally, due to the limitations of pose detection techniques, the extracted pose sequences can sometimes be inaccurate, leading to low-quality video outputs. We introduce a novel task aimed at generating human motion videos solely from reference images and natural language. This approach offers greater flexibility and ease of use, as text is more accessible than the desired guidance videos. However, training an end-to-end model for this task requires millions of high-quality text and human motion video pairs, which are challenging to obtain. To address this, we propose a new framework called Fleximo, which leverages large-scale pre-trained text-to-3D motion models. This approach is not straightforward, as the text-generated skeletons may not consistently match the scale of the reference image and may lack detailed information. To overcome these challenges, we introduce an anchor point based rescale method and design a skeleton adapter to fill in missing details and bridge the gap between text-to-motion and motion-to-video generation. We also propose a video refinement process to further enhance video quality. A large language model (LLM) is employed to decompose natural language into discrete motion sequences, enabling the generation of motion videos of any desired length. To assess the performance of Fleximo, we introduce a new benchmark called MotionBench, which includes 400 videos across 20 identities and 20 motions. We also propose a new metric, MotionScore, to evaluate the accuracy of motion following. Both qualitative and quantitative results demonstrate that our method outperforms existing text-conditioned image-to-video generation methods. All code and model weights will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2411.19459

Code (0)

등록된 구현이 없습니다.

Tasks

Image to Video GenerationLarge Language ModelText to 3DVideo Generation

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

FlexiMo: A Flexible Remote Sensing Foundation Model

2025-03-31 · Xuyang Li, Chenyu Li, Pedram Ghamisi, Danfeng Hong

The rapid expansion of multi-source satellite imagery drives innovation in Earth observation, opening unprecedented opportunities for Remote Sensing Foundation Models to harness diverse data. However, many existing model…

Cloud DetectionEarth ObservationLand Cover Classificationmodel+1

Text2Performer: Text-Driven Human Video Generation

2023-04-17 · ICCV 2023 1 · Yuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu 외

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts des…

Video Generation

XmoPipe: A Pipeline for Large-Scale In-the-Wild Human Motion Dataset Construction

2026-06-17 · Nathan Salazar, Emmanuel Dellandréa, Mathieu Lefort, Alexandre Meyer arxiv

Large-scale human motion datasets are essential for training robust motion models for analysis, synthesis, and understanding. While marker-based motion capture provides precise data, it is costly and limited in scale and…

Unimotion: Unifying 3D Human Motion Synthesis and Understanding

2024-09-24 · Chuqiao Li, Julian Chibane, Yannan He, Naama Pearl 외

We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control avatar motion with global text conditioni…

Motion Synthesis

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

2025-08-12 · Jingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao 외 arxiv

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground s…

Video Generation