paper-with-me

홈 › Papers

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

2025-07-31 · Shiyao Yu, Zi-An Wang, Kangning Yin, Zheng Tian, Mingyuan Zhang, Weixin Si, Shihao Zou arxiv

Motion retrieval is crucial for motion acquisition, offering superior precision, realism, controllability, and editability compared to motion generation. Existing approaches leverage contrastive learning to construct a unified embedding space for motion retrieval from text or visual modality. However, these methods lack a more intuitive and user-friendly interaction mode and often overlook the sequential representation of most modalities for improved retrieval performance. To address these limitations, we propose a framework that aligns four modalities -- text, audio, video, and motion -- within a fine-grained joint embedding space, incorporating audio for the first time in motion retrieval to enhance user immersion and convenience. This fine-grained space is achieved through a sequence-level contrastive learning approach, which captures critical details across modalities for better alignment. To evaluate our framework, we augment existing text-motion datasets with synthetic but diverse audio recordings, creating two multi-modal motion retrieval datasets. Experimental results demonstrate superior performance over state-of-the-art methods across multiple sub-tasks, including an 10.16% improvement in R@10 for text-to-motion retrieval and a 25.43% improvement in R@1 for video-to-motion retrieval on the HumanML3D dataset. Furthermore, our results show that our 4-modal framework significantly outperforms its 3-modal counterpart, underscoring the potential of multi-modal motion retrieval for advancing motion acquisition.

📄 PDF Abstract BibTeX arXiv:2507.23188

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning

2026-01-29 · Hanmo Chen, Guangtao Lyu, Chenghao Xu, Jiexi Yan 외 arxiv

As a foundational task in human-centric cross-modal intelligence, motion-language retrieval aims to bridge the semantic gap between natural language and human motion, enabling intuitive motion analysis, yet existing appr…

Fine-Grained Instance-Level Sketch-Based Video Retrieval

2020-02-21 · Peng Xu, Kun Liu, Tao Xiang, Timothy M. Hospedales 외

Existing sketch-analysis work studies sketches depicting static objects or scenes. In this work, we propose a novel cross-modal retrieval problem of fine-grained instance-level sketch-based video retrieval (FG-SBVR), whe…

Cross-Modal RetrievalImage RetrievalRetrievalVideo Retrieval

Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction

2026-03-10 · Yao Zhang, Zhuchenyang Liu, Yanlan He, Thomas Ploetz 외 arxiv

Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search across the two modalities. Most existing…

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

2026-07-30 · Bohan Hou, Haoqiang Lin, Xuemeng Song, Haokun Wen 외 arxiv

Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing d…

Image RetrievalVisual Dialog

Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition

2026-05-16 · Zeheng Wang, Bo Zhao, Yijie Zhu, Zhishu Liu 외 arxiv

Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal reasoning, they typically treat emotion …

Multimodal Emotion RecognitionEmotion ClassificationMultimodal Reasoning