paper-with-me

홈 › Papers

MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction

2025-07-09 · Yin Wang, Mu li, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang arxiv

We introduce MOST, a novel motion diffusion model via temporal clip Banzhaf interaction, aimed at addressing the persistent challenge of generating human motion from rare language prompts. While previous approaches struggle with coarse-grained matching and overlook important semantic cues due to motion redundancy, our key insight lies in leveraging fine-grained clip relationships to mitigate these issues. MOST's retrieval stage presents the first formulation of its kind - temporal clip Banzhaf interaction - which precisely quantifies textual-motion coherence at the clip level. This facilitates direct, fine-grained text-to-motion clip matching and eliminates prevalent redundancy. In the generation stage, a motion prompt module effectively utilizes retrieved motion clips to produce semantically consistent movements. Extensive evaluations confirm that MOST achieves state-of-the-art text-to-motion retrieval and generation performance by comprehensively addressing previous challenges, as demonstrated through quantitative and qualitative results highlighting its effectiveness, especially for rare prompts.

📄 PDF Abstract BibTeX arXiv:2507.06590

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DRFusion: Drift-Resilient Temporally Consistent Infrared-Visible Video Fusion

2026-05-25 · Xingyuan Li, Haoyuan Xu, Shulin Li, Xiang Chen 외 arxiv

Infrared and visible video fusion is essential for achieving comprehensive perception in dynamic scenes. However, maintaining temporal consistency remains a formidable challenge. Conventional methods relying on optical f…

MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching

2025-02-18 · Yen-Siang Wu, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang

Text-to-video (T2V) diffusion models have shown promising capabilities in synthesizing realistic videos from input text prompts. However, the input text description alone provides limited control over the precise objects…

MoVideo: Motion-Aware Video Generation with Diffusion Models

2023-11-19 · Jingyun Liang, Yuchen Fan, Kai Zhang, Radu Timofte 외

While recent years have witnessed great progress on using diffusion models for video generation, most of them are simple extensions of image generation frameworks, which fail to explicitly consider one of the key differe…

Image GenerationImage to Video GenerationOptical Flow EstimationVideo Generation

Text-driven Human Motion Generation with Motion Masked Diffusion Model

2024-09-29 · Xingyu Chen

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, wh…

DiversityMotion Generation

RecMoDiffuse: Recurrent Flow Diffusion for Human Motion Generation

2024-06-11 · Mirgahney Mohamed, Harry Jake Cunningham, Marc P. Deisenroth, Lourdes Agapito

Human motion generation has paramount importance in computer animation. It is a challenging generative temporal modelling task due to the vast possibilities of human motion, high human sensitivity to motion coherence and…

Motion Generation