paper-with-me

홈 › Papers

MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations

2024-10-17 · Liang Xu, Shaoyang Hua, Zili Lin, Yifan Liu, Feipeng Ma, Yichao Yan, Xin Jin, Xiaokang Yang, Wenjun Zeng

In this paper, we tackle the problem of how to build and benchmark a large motion model (LMM). The ultimate goal of LMM is to serve as a foundation model for versatile motion-related tasks, e.g., human motion generation, with interpretability and generalizability. Though advanced, recent LMM-related works are still limited by small-scale motion data and costly text descriptions. Besides, previous motion benchmarks primarily focus on pure body movements, neglecting the ubiquitous motions in context, i.e., humans interacting with humans, objects, and scenes. To address these limitations, we consolidate large-scale video action datasets as knowledge banks to build MotionBank, which comprises 13 video action datasets, 1.24M motion sequences, and 132.9M frames of natural and diverse human motions. Different from laboratory-captured motions, in-the-wild human-centric videos contain abundant motions in context. To facilitate better motion text alignment, we also meticulously devise a motion caption generation algorithm to automatically produce rule-based, unbiased, and disentangled text descriptions via the kinematic characteristics for each motion. Extensive experiments show that our MotionBank is beneficial for general motion-related tasks of human motion generation, motion in-context generation, and motion understanding. Video motions together with the rule-based text annotations could serve as an efficient alternative for larger LMMs. Our dataset, codes, and benchmark will be publicly available at https://github.com/liangxuy/MotionBank.

📄 PDF Abstract BibTeX arXiv:2410.13790

Code (1)

liangxuy/motionbank 공식 구현

Tasks

Caption GenerationMotion Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

2023-08-16 · ICCV 2023 1 · Henghui Ding, Chang Liu, Shuting He, Xudong Jiang 외

This paper strives for motion expressions guided video segmentation, which focuses on segmenting objects in video content based on a sentence describing the motion of the objects. Existing referring video object datasets…

Motion Expressions Guided Video SegmentationObjectReferring Video Object SegmentationSegmentation+5

FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos

2025-12-11 · Yulu Gan, Ligeng Zhu, Dandan Shan, Baifeng Shi 외 arxiv

Motion understanding is fundamental to physical reasoning, enabling models to infer dynamics and predict future states. However, state-of-the-art models still struggle on recent motion benchmarks, primarily due to the sc…

Spatial Reasoning

Fleximo: Towards Flexible Text-to-Human Motion Video Generation

2024-11-29 · Yuhang Zhang, Yuan Zhou, Zeyu Liu, Yuxuan Cai 외

Current methods for generating human motion videos rely on extracting pose sequences from reference videos, which restricts flexibility and control. Additionally, due to the limitations of pose detection techniques, the …

Image to Video GenerationLarge Language ModelText to 3DVideo Generation

Towards Understanding Camera Motions in Any Video

2025-04-21 · Zhiqiu Lin, Siyuan Cen, Daniel Jiang, Jay Karhade 외

We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous mu…

Question AnsweringText RetrievalVideo Question AnsweringVideo-Text Retrieval

H-VFI: Hierarchical Frame Interpolation for Videos with Large Motions

2022-11-21 · Changlin Li, Guangyang Wu, Yanan sun, Xin Tao 외

Capitalizing on the rapid development of neural networks, recent video frame interpolation (VFI) methods have achieved notable improvements. However, they still fall short for real-world videos containing large motions. …

Video Frame Interpolation