paper-with-me

홈 › Papers

FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing

2025-07-26 · Bizhu Wu, Jinheng Xie, Meidan Ding, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, Linlin Shen arxiv

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching the textual description with more details. Specifically, we propose the FineMotion dataset, which contains over 442,000 human motion snippets - short segments of human motion sequences - and their corresponding detailed descriptions of human body part movements. Additionally, the dataset includes about 95k detailed paragraphs describing the movements of human body parts of entire motion sequences. Experimental results demonstrate the significance of our dataset on the text-driven finegrained human motion generation task, especially with a remarkable +15.3% improvement in Top-3 accuracy for the MDM model. Notably, we further support a zero-shot pipeline of fine-grained motion editing, which focuses on detailed editing in both spatial and temporal dimensions via text. Dataset and code available at: CVI-SZU/FineMotion

📄 PDF Abstract BibTeX arXiv:2507.19850

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Motion Generation from Fine-grained Textual Descriptions

2024-03-20 · Kunhang Li, Yansong Feng

The task of text2motion is to generate human motion sequences from given textual descriptions, where the model explores diverse mappings from natural language instructions to human body movements. While most existing wor…

Motion Generation

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

2025-01-14 · CVPR 2025 1 · Hongyu Li, Jinyu Chen, Ziyu Wei, Shaofei Huang 외

Recent advancements in multimodal large language models (MLLMs) have shown promising results, yet existing approaches struggle to effectively handle both temporal and spatial localization simultaneously. This challenge s…

Feature CompressionLanguage ModelingLanguage ModellingLarge Language Model+3

Spatio-Temporal Analysis of Facial Actions using Lifecycle-Aware Capsule Networks

2020-11-17 · Nikhil Churamani, Sinan Kalkan, Hatice Gunes

Most state-of-the-art approaches for Facial Action Unit (AU) detection rely upon evaluating facial expressions from static frames, encoding a snapshot of heightened facial activity. In real-world interactions, however, f…

Spatially Encoding Temporal Correlations to Classify Temporal Data Using Convolutional Neural Networks

2015-09-24 · Zhiguang Wang, Tim Oates

We propose an off-line approach to explicitly encode temporal patterns spatially as different types of images, namely, Gramian Angular Fields and Markov Transition Fields. This enables the use of techniques from computer…

ClassificationGeneral ClassificationTime SeriesTime Series Analysis

SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

2025-03-18 · Jiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li 외

Multimodal large language models (MLLMs) have made remarkable progress in either temporal or spatial localization. However, they struggle to perform spatio-temporal video grounding. This limitation stems from two major c…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3