paper-with-me

홈 › Papers

MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation

2025-05-16 · Gabriel Maldonado, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Vinit Katariya, Hamed Tabkhi

Human motion generation is essential for fields such as animation, robotics, and virtual reality, requiring models that effectively capture motion dynamics from text descriptions. Existing approaches often rely on Contrastive Language-Image Pretraining (CLIP)-based text encoders, but their training on text-image pairs constrains their ability to understand temporal and kinematic structures inherent in motion and motion generation. This work introduces MoCLIP, a fine-tuned CLIP model with an additional motion encoding head, trained on motion sequences using contrastive learning and tethering loss. By explicitly incorporating motion-aware representations, MoCLIP enhances motion fidelity while remaining compatible with existing CLIP-based pipelines and seamlessly integrating into various CLIP-based methods. Experiments demonstrate that MoCLIP improves Top-1, Top-2, and Top-3 accuracy while maintaining competitive FID, leading to improved text-to-motion alignment results. These results highlight MoCLIP's versatility and effectiveness, establishing it as a robust framework for enhancing motion generation.

📄 PDF Abstract BibTeX arXiv:2505.10810

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningMotion Generation

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors

2025-09-21 · Binhua Huang, Ni Wang, Arjun Pakrashi, Soumyabrata Dev arxiv

Video action recognition is a fundamental task in computer vision, but state-of-the-art models are often computationally expensive and rely on extensive video pre-training. In parallel, large-scale vision-language models…

Action Recognition

CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging

2024-07-10 · Raza Imam, Mohammed Talha Alam, Umaima Rahman, Mohsen Guizani 외

Existing vision-text contrastive learning models enhance representation transferability and support zero-shot prediction by matching paired image and caption embeddings while pushing unrelated pairs apart. However, astro…

Contrastive LearningImage-text RetrievalRetrievalText Retrieval+2

HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models

2025-11-27 · Haoxi Zeng, Haoxuan Li, Yi Bin, Pengpeng Zeng 외 arxiv

Contrastive Language-Image Pre-training (CLIP) has demonstrated remarkable generalization ability and strong performance across a wide range of vision-language tasks. However, due to the lack of region-level supervision,…

AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation

2025-08-31 · Gyehun Go, Satbyul Han, Ahyeon Choi, Eunjin Choi 외 arxiv

Recent advances in text-to-music (TTM) generation have enabled controllable and expressive music creation using natural language prompts. However, the emotional fidelity of TTM systems remains largely underexplored compa…

Text-to-Music Generation

MOCLIP: A Foundation Model for Large-Scale Nanophotonic Inverse Design

2025-11-24 · S. Rodionov, A. Burguete-Lopez, M. Makarenko, Q. Wang 외 arxiv

Foundation models (FM) are transforming artificial intelligence by enabling generalizable, data-efficient solutions across different domains for a broad range of applications. However, the lack of large and diverse datas…

Contrastive Learning