paper-with-me

홈 › Papers

Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models

2024-07-22 · Kent Fujiwara, Mikihiro Tanaka, Qing Yu

With the release of large-scale motion datasets with textual annotations, the task of establishing a robust latent space for language and 3D human motion has recently witnessed a surge of interest. Methods have been proposed to convert human motion and texts into features to achieve accurate correspondence between them. Despite these efforts to align language and motion representations, we claim that the temporal element is often overlooked, especially for compound actions, resulting in chronological inaccuracies. To shed light on the temporal alignment in motion-language latent spaces, we propose Chronologically Accurate Retrieval (CAR) to evaluate the chronological understanding of the models. We decompose textual descriptions into events, and prepare negative text samples by shuffling the order of events in compound action descriptions. We then design a simple task for motion-language models to retrieve the more likely text from the ground truth and its chronologically shuffled version. CAR reveals many cases where current motion-language models fail to distinguish the event chronology of human motion, despite their impressive performance in terms of conventional evaluation metrics. To achieve better temporal alignment between text and motion, we further propose to use these texts with shuffled sequence of events as negative samples during training to reinforce the motion-language models. We conduct experiments on text-motion retrieval and text-to-motion generation using the reinforced motion-language models, which demonstrate improved performance over conventional approaches, indicating the necessity to consider temporal elements in motion-language alignment.

📄 PDF Abstract BibTeX arXiv:2407.15408

Code (0)

등록된 구현이 없습니다.

Tasks

Motion GenerationRetrieval

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

MoChat: Joints-Grouped Spatio-Temporal Grounding LLM for Multi-Turn Motion Comprehension and Description

2024-10-15 · Jiawei Mo, Yixuan Chen, Rifen Lin, Yongkang Ni 외

Despite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interac…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

2024-11-21 · Weiheng Lu, Jian Li, An Yu, Ming-Ching Chang 외

Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain challenging due to LLMs' limited context si…

Moment RetrievalNatural Language Moment RetrievalRetrieval

T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding

2026-03-07 · Chaohong Guo, Yihan He, Yongwei Nie, Fei Ma 외 arxiv

Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dynamics. Existing Vision-LMMs typically pe…

Temporal Sequences

LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval

2025-10-07 · Avishree Khare, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos 외 arxiv

Neural models such as YOLO and HuBERT can be used to detect local properties such as objects ("car") and emotions ("angry") in individual frames of videos and audio clips respectively. The likelihood of these detections …

Video Retrieval

A Survey on Event-driven 3D Reconstruction: Development under Different Categories

2025-03-25 · Chuanzhi Xu, Haoxian Zhou, Haodong Chen, Vera Chung 외

Event cameras have gained increasing attention for 3D reconstruction due to their high temporal resolution, low latency, and high dynamic range. They capture per-pixel brightness changes asynchronously, allowing accurate…

3D Reconstruction