paper-with-me

홈 › Papers

Exploring Motion-Language Alignment for Text-driven Motion Generation

2026-04-03 · Ruxi Gu, Zilei Wang, Wei Wang arxiv

Text-driven human motion generation aims to synthesize realistic motion sequences that follow textual descriptions. Despite recent advances, accurately aligning motion dynamics with textual semantics remains a fundamental challenge. In this paper, we revisit text-to-motion generation from the perspective of motion-language alignment and propose MLA-Gen, a framework that integrates global motion priors with fine-grained local conditioning. This design enables the model to capture common motion patterns, while establishing detailed alignment between texts and motions. Furthermore, we identify a previously overlooked attention sink phenomenon in human motion generation, where attention disproportionately concentrates on the start text token, limiting the utilization of informative textual cues and leading to degraded semantic grounding. To analyze this issue, we introduce SinkRatio, a metric for measuring attention concentration, and develop alignment-aware masking and control strategies to regulate attention during generation. Extensive experiments demonstrate that our approach consistently improves both motion quality and motion-language alignment over strong baselines. Code will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2604.02973

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition

2024-05-31 · Weichao Zhao, Hezhen Hu, Wengang Zhou, Yunyao Mao 외

Sign language recognition (SLR) has long been plagued by insufficient model representation capabilities. Although current pre-training approaches have alleviated this dilemma to some extent and yielded promising performa…

Self-Supervised LearningSign Language Recognition

JeSemE: Interleaving Semantics and Emotions in a Web Service for the Exploration of Language Change Phenomena

2018-08-01 · COLING 2018 8 · Johannes Hellrich, Sven Buechel, Udo Hahn

We here introduce a substantially extended version of JeSemE, an interactive website for visually exploring computationally derived time-variant information on word meanings and lexical emotions assembled from five large…

Sentiment AnalysisWord Embeddings

JeSemE: A Website for Exploring Diachronic Changes in Word Meaning and Emotion

2018-07-11 · Johannes Hellrich, Sven Buechel, Udo Hahn

We here introduce a substantially extended version of JeSemE, an interactive website for visually exploring computationally derived time-variant information on word meanings and lexical emotions assembled from five large…

RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control

2025-06-15 · Junpeng Yue, Zepeng Wang, Yuxuan Wang, Weishuai Zeng 외

This paper focuses on a critical challenge in robotics: translating text-driven human motions into executable actions for humanoid robots, enabling efficient and cost-effective learning of new behaviors. While existing t…

Humanoid ControlMotion GenerationSemantic correspondence

HumanTOMATO: Text-aligned Whole-body Motion Generation

2023-10-19 · Shunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin 외

This work targets a novel text-driven whole-body motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and …

Motion GenerationMotion Synthesis