paper-with-me

홈 › Papers

Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning

2026-01-29 · Hanmo Chen, Guangtao Lyu, Chenghao Xu, Jiexi Yan, Xu Yang, Cheng Deng arxiv

As a foundational task in human-centric cross-modal intelligence, motion-language retrieval aims to bridge the semantic gap between natural language and human motion, enabling intuitive motion analysis, yet existing approaches predominantly focus on aligning entire motion sequences with global textual representations. This global-centric paradigm overlooks fine-grained interactions between local motion segments and individual body joints and text tokens, inevitably leading to suboptimal retrieval performance. To address this limitation, we draw inspiration from the pyramidal process of human motion perception (from joint dynamics to segment coherence, and finally to holistic comprehension) and propose a novel Pyramidal Shapley-Taylor (PST) learning framework for fine-grained motion-language retrieval. Specifically, the framework decomposes human motion into temporal segments and spatial body joints, and learns cross-modal correspondences through progressive joint-wise and segment-wise alignment in a pyramidal fashion, effectively capturing both local semantic details and hierarchical structural relationships. Extensive experiments on multiple public benchmark datasets demonstrate that our approach significantly outperforms state-of-the-art methods, achieving precise alignment between motion segments and body joints and their corresponding text tokens. The code of this work will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2601.21904

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KinMo: Kinematic-aware Human Motion Understanding and Generation

2024-11-23 · Pengfei Zhang, Pinxin Liu, Hyeongwoo Kim, Pablo Garrido 외

Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as ``run", fails …

Motion GenerationMotion Synthesis

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

2025-09-20 · Sirui Wang, Andong Chen, Tiejun Zhao arxiv

Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predefined labels, reference audio, or natural…

Speech Synthesis

Micro-AU CLIP: Fine-Grained Contrastive Learning from Local Independence to Global Dependency for Micro-Expression Action Unit Detection

2026-03-17 · Jinsheng Wei, Fengzhou Guo, Yante Li, Haoyu Chen 외 arxiv

Micro-expression (ME) action units (Micro-AUs) provide objective clues for fine-grained genuine emotion analysis. Most existing Micro-AU detection methods learn AU features from the whole facial image/video, which confli…

Action Unit DetectionContrastive Learning

MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment

2026-02-22 · Duc Duy Nguyen, Tat-Jun Chin, Minh Hoai arxiv

We aim to learn a joint representation between inertial measurement unit (IMU) signals and 2D pose sequences extracted from video, enabling accurate cross-modal retrieval, temporal synchronization, subject and body-part …

Cross-Modal RetrievalContrastive LearningAction Recognition

Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction

2026-03-10 · Yao Zhang, Zhuchenyang Liu, Yanlan He, Thomas Ploetz 외 arxiv

Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search across the two modalities. Most existing…