paper-with-me

홈 › Papers

Hierarchical Motion Captioning Utilizing External Text Data Source

2025-09-01 · Clayton Leite, Yu Xiao arxiv

This paper introduces a novel approach to enhance existing motion captioning methods, which directly map representations of movement to high-level descriptive captions (e.g., `a person doing jumping jacks"). The existing methods require motion data annotated with high-level descriptions (e.g., jumping jacks"). However, such data is rarely available in existing motion-text datasets, which additionally do not include low-level motion descriptions. To address this, we propose a two-step hierarchical approach. First, we employ large language models to create detailed descriptions corresponding to each high-level caption that appears in the motion-text datasets (e.g., jumping while synchronizing arm extensions with the opening and closing of legs" for `jumping jacks"). These refined annotations are used to retrain motion-to-text models to produce captions with low-level details. Second, we introduce a pioneering retrieval-based mechanism. It aligns the detailed low-level captions with candidate high-level captions from additional text data sources, and combine them with motion features to fabricate precise high-level captions. Our methodology is distinctive in its ability to harness knowledge from external text sources to greatly increase motion captioning accuracy, especially for movements not covered in existing motion-text datasets. Experiments on three distinct motion-text datasets (HumanML3D, KIT, and BOTH57M) demonstrate that our method achieves an improvement in average performance (across BLEU-1, BLEU-4, CIDEr, and ROUGE-L) ranging from 6% to 50% compared to the state-of-the-art M2T-Interpretable.

📄 PDF Abstract BibTeX arXiv:2509.01471

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Captioning

Similar Papers 제목 키워드 기반

HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning

2024-12-19 · Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 외

With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several stu…

Dense Video CaptioningVideo Captioning

Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning

2026-06-17 · Minh-Loi Nguyen, Xuan-Vu Le, Long-Bao Nguyen, Hoang-Bach Ngo 외 arxiv

Traditional image captioning methods often struggle to generate comprehensive, context-rich descriptions, especially for details not directly observable from visual cues. To overcome this, we propose a novel retrieval-au…

Image Captioning

Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning

2018-11-07 · Xin Wang, Jiawei Wu, Da Zhang, Yu Su 외

Although promising results have been achieved in video captioning, existing models are limited to the fixed inventory of activities in the training corpus, and do not generalize to open vocabulary scenarios. Here we intr…

Mixture-of-ExpertsVideo Captioning

Motion Guided Region Message Passing for Video Captioning

2021-01-01 · ICCV 2021 10 · Shaoxiang Chen, Yu-Gang Jiang

Video captioning is an important vision task and has been intensively studied in the computer vision community. Existing methods that utilize the fine-grained spatial information have achieved significant improvement…

DecoderVideo Captioning

Spatial Attention as an Interface for Image Captioning Models

2020-09-29 · Philipp Sadler

The internal workings of modern deep learning models stay often unclear to an external observer, although spatial attention mechanisms are involved. The idea of this work is to translate these spatial attentions into nat…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)