paper-with-me

홈 › Papers

MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization

2024-05-06 · Massimiliano Pappa, Luca Collorone, Giovanni Ficarra, Indro Spinelli, Fabio Galasso

Diffusion Models have revolutionized the field of human motion generation by offering exceptional generation quality and fine-grained controllability through natural language conditioning. Their inherent stochasticity, that is the ability to generate various outputs from a single input, is key to their success. However, this diversity should not be unrestricted, as it may lead to unlikely generations. Instead, it should be confined within the boundaries of text-aligned and realistic generations. To address this issue, we propose MoDiPO (Motion Diffusion DPO), a novel methodology that leverages Direct Preference Optimization (DPO) to align text-to-motion models. We streamline the laborious and expensive process of gathering human preferences needed in DPO by leveraging AI feedback instead. This enables us to experiment with novel DPO strategies, using both online and offline generated motion-preference pairs. To foster future research we contribute with a motion-preference dataset which we dub Pick-a-Move. We demonstrate, both qualitatively and quantitatively, that our proposed method yields significantly more realistic motions. In particular, MoDiPO substantially improves Frechet Inception Distance (FID) while retaining the same RPrecision and Multi-Modality performances.

📄 PDF Abstract BibTeX arXiv:2405.03803

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityMotion Generation

Methods 이 논문이 사용한 방법론

DPO 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization

2024-12-06 · Xiaofeng Tan, Hongsong Wang, Xin Geng, Pan Zhou

Text-to-motion generation is essential for advancing the creative industry but often presents challenges in producing consistent, realistic motions. To address this, we focus on fine-tuning text-to-motion models to consi…

Motion Generation

RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control

2025-06-15 · Junpeng Yue, Zepeng Wang, Yuxuan Wang, Weishuai Zeng 외

This paper focuses on a critical challenge in robotics: translating text-driven human motions into executable actions for humanoid robots, enabling efficient and cost-effective learning of new behaviors. While existing t…

Humanoid ControlMotion GenerationSemantic correspondence

Exploring Motion-Language Alignment for Text-driven Motion Generation

2026-04-03 · Ruxi Gu, Zilei Wang, Wei Wang arxiv

Text-driven human motion generation aims to synthesize realistic motion sequences that follow textual descriptions. Despite recent advances, accurately aligning motion dynamics with textual semantics remains a fundamenta…

EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

2025-11-25 · Jingyang Jia, Kai Shu, Gang Yang, Long Xing 외 arxiv

Continuous emotional image content generation (C-EICG) is emerging rapidly due to its ability to produce images aligned with both user descriptions and continuous emotional values. However, existing approaches lack emoti…

Image Generation

HumanTOMATO: Text-aligned Whole-body Motion Generation

2023-10-19 · Shunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin 외

This work targets a novel text-driven whole-body motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and …

Motion GenerationMotion Synthesis