Hierarchical Motion Captioning Utilizing External Text Data Source
This paper introduces a novel approach to enhance existing motion captioning methods, which directly map representations of movement to high-level descriptive captions (e.g., `a person doing jumping jacks"). The existing methods require motion data annotated with high-level descriptions (e.g., jumping jacks"). However, such data is rarely available in existing motion-text datasets, which additionally do not include low-level motion descriptions. To address this, we propose a two-step hierarchical approach. First, we employ large language models to create detailed descriptions corresponding to each high-level caption that appears in the motion-text datasets (e.g., jumping while synchronizing arm extensions with the opening and closing of legs" for `jumping jacks"). These refined annotations are used to retrain motion-to-text models to produce captions with low-level details. Second, we introduce a pioneering retrieval-based mechanism. It aligns the detailed low-level captions with candidate high-level captions from additional text data sources, and combine them with motion features to fabricate precise high-level captions. Our methodology is distinctive in its ability to harness knowledge from external text sources to greatly increase motion captioning accuracy, especially for movements not covered in existing motion-text datasets. Experiments on three distinct motion-text datasets (HumanML3D, KIT, and BOTH57M) demonstrate that our method achieves an improvement in average performance (across BLEU-1, BLEU-4, CIDEr, and ROUGE-L) ranging from 6% to 50% compared to the state-of-the-art M2T-Interpretable.
Code (0)
등록된 구현이 없습니다.
Tasks
Motion CaptioningSimilar Papers 제목 키워드 기반
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several stu…
Dense Video CaptioningVideo CaptioningHierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning
Traditional image captioning methods often struggle to generate comprehensive, context-rich descriptions, especially for details not directly observable from visual cues. To overcome this, we propose a novel retrieval-au…
Image CaptioningLearning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning
Although promising results have been achieved in video captioning, existing models are limited to the fixed inventory of activities in the training corpus, and do not generalize to open vocabulary scenarios. Here we intr…
Mixture-of-ExpertsVideo CaptioningMotion Guided Region Message Passing for Video Captioning
Video captioning is an important vision task and has been intensively studied in the computer vision community. Existing methods that utilize the fine-grained spatial information have achieved significant improvement…
DecoderVideo CaptioningSpatial Attention as an Interface for Image Captioning Models
The internal workings of modern deep learning models stay often unclear to an external observer, although spatial attention mechanisms are involved. The idea of this work is to translate these spatial attentions into nat…
Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)