paper-with-me

홈 › Papers

ViMo: Generating Motions from Casual Videos

2024-08-13 · Liangdong Qiu, Chengxing Yu, Yanran Li, Zhao Wang, Haibin Huang, Chongyang Ma, Di Zhang, Pengfei Wan, Xiaoguang Han

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion generation methods predominantly rely on manually collected motion datasets, usually tediously sourced from motion capture (Mocap) systems or Multi-View cameras, unavoidably resulting in a limited size that severely undermines their generalizability. Inspired by recent advance of diffusion models, we probe a simple and effective way to capture motions from videos and propose a novel Video-to-Motion-Generation framework (ViMo) which could leverage the immense trove of untapped video content to produce abundant and diverse 3D human motions. Distinct from prior work, our videos could be more causal, including complicated camera movements and occlusions. Striking experimental results demonstrate the proposed model could generate natural motions even for videos where rapid movements, varying perspectives, or frequent occlusions might exist. We also show this work could enable three important downstream applications, such as generating dancing motions according to arbitrary music and source video style. Extensive experimental results prove that our model offers an effective and scalable way to generate diversity and realistic motions. Code and demos will be public soon.

📄 PDF Abstract BibTeX arXiv:2408.06614

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

EVIMO2: An Event Camera Dataset for Motion Segmentation, Optical Flow, Structure from Motion, and Visual Inertial Odometry in Indoor Scenes with Monocular or Stereo Algorithms

2022-05-06 · Levi Burner, Anton Mitrokhin, Cornelia Fermüller, Yiannis Aloimonos

A new event camera dataset, EVIMO2, is introduced that improves on the popular EVIMO dataset by providing more data, from better cameras, in more complex scenarios. As with its predecessor, EVIMO2 provides labels in the …

Motion SegmentationOptical Flow EstimationSegmentationSemantic Segmentation

TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos

2024-03-26 · Yufu Wang, ZiYun Wang, Lingjie Liu, Kostas Daniilidis

We propose TRAM, a two-stage method to reconstruct a human's global trajectory and motion from in-the-wild videos. TRAM robustifies SLAM to recover the camera motion in the presence of dynamic humans and uses the scene b…

3D Human Pose Estimation

Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos

2024-08-01 · Subin Jeon, In Cho, Minsu Kim, Woong Oh Cho 외

We propose a new framework for creating and easily manipulating 3D models of arbitrary objects using casually captured videos. Our core ingredient is a novel hierarchy deformation model, which captures motions of objects…

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

2025-10-30 · Jing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang 외 arxiv

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generativ…

Video Generation

ViMo: A Generative Visual GUI World Model for App Agents

2025-04-15 · Dezhao Luo, Bohan Tang, Kang Li, Georgios Papoudakis 외

App agents, which autonomously operate mobile Apps through Graphical User Interfaces (GUIs), have gained significant interest in real-world applications. Yet, they often struggle with long-horizon planning, failing to fi…