paper-with-me

Papers

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

2025-10-30 · Jing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang, Guorui Song, Ailing Zeng, Xian Liu, Chen Wei, Wanqi Yin, Qingping Sun, Zhongang Cai, Lei Yang, Ziwei Liu arxiv

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most notably video generation (ViGen), have demonstrated remarkable generalization in modeling human behaviors, highlighting transferable insights that MoGen can leverage. Motivated by this observation, we present a comprehensive framework that systematically transfers knowledge from ViGen to MoGen across three key pillars: data, modeling, and evaluation. First, we introduce ViMoGen-228K, a large-scale dataset comprising 228,000 high-quality motion samples that integrates high-fidelity optical MoCap data with semantically annotated motions from web videos and synthesized samples generated by state-of-the-art ViGen models. The dataset includes both text-motion pairs and text-video-motion triplets, substantially expanding semantic diversity. Second, we propose ViMoGen, a flow-matching-based diffusion transformer that unifies priors from MoCap data and ViGen models through gated multimodal conditioning. To enhance efficiency, we further develop ViMoGen-light, a distilled variant that eliminates video generation dependencies while preserving strong generalization. Finally, we present MBench, a hierarchical benchmark designed for fine-grained evaluation across motion quality, prompt fidelity, and generalization ability. Extensive experiments show that our framework significantly outperforms existing approaches in both automatic and human evaluations. The code, data, and benchmark will be made publicly available. Homepage: https://motrixlab.github.io/2026_iclr_vimogen.

📄 PDF Abstract BibTeX arXiv:2510.26794

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

2024-12-15 · CVPR 2025 1 · Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Pedro M B Rezende 외

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object …

Autonomous Driving

Data-driven Head Motion Generation through Natural Gaze-Head Coordination

2026-05-25 · Xiaohan Liu, Yilin Wen, Yusuke Sugano arxiv

We present the first data-driven approach to model temporal gaze-head coordination from large-scale in-the-wild facial videos. To obtain training data for generalizable learning, we propose an automatic pipeline that ext…

Video Generation

Learning Generalizable Human Motion Generator with Reinforcement Learning

2024-05-24 · Yunyao Mao, Xiaoyang Liu, Wengang Zhou, Zhenbo Lu 외

Text-driven human motion generation, as one of the vital tasks in computer-aided content creation, has recently attracted increasing attention. While pioneering research has largely focused on improving numerical perform…

Motion Generationreinforcement-learningReinforcement Learning

LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization

2026-04-28 · Huyen Nguyen, Haoxuan Zhang, Yang Zhang, Junhua Ding 외 arxiv

Evaluating long document summaries remains the primary bottleneck in summarization research. Existing metrics correlate weakly with human judgments and produce aggregate scores without explaining deficiencies or guiding …

Document SummarizationText Generation

Persona-Based Synthetic Data Generation Using Multi-Stage Conditioning with Large Language Models for Emotion Recognition

2025-07-15 · Keito Inoshita, Rushia Harada arxiv

In the field of emotion recognition, the development of high-performance models remains a challenge due to the scarcity of high-quality, diverse emotional datasets. Emotional expressions are inherently subjective, shaped…

Synthetic Data GenerationEmotion ClassificationEmotion Recognition