paper-with-me

Papers

Controllable Video Generation through Global and Local Motion Dynamics

2022-04-13 · Aram Davtyan, Paolo Favaro

We present GLASS, a method for Global and Local Action-driven Sequence Synthesis. GLASS is a generative model that is trained on video sequences in an unsupervised manner and that can animate an input image at test time. The method learns to segment frames into foreground-background layers and to generate transitions of the foregrounds over time through a global and local action representation. Global actions are explicitly related to 2D shifts, while local actions are instead related to (both geometric and photometric) local deformations. GLASS uses a recurrent neural network to transition between frames and is trained through a reconstruction loss. We also introduce W-Sprites (Walking Sprites), a novel synthetic dataset with a predefined action space. We evaluate our method on both W-Sprites and real datasets, and find that GLASS is able to generate realistic video sequences from a single input image and to successfully learn a more advanced action space than in prior work.

📄 PDF Abstract BibTeX arXiv:2204.06558

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories

2026-02-16 · Zun Wang, Han Lin, Jaehong Yoon, Jaemin Cho 외 arxiv

Maintaining spatial world consistency over long horizons remains a central challenge for camera-controllable video generation. Existing memory-based approaches often condition generation on globally reconstructed 3D scen…

Depth EstimationVideo Generation

EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

2025-11-22 · Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy 외 arxiv

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controll…

Video GenerationVideo Prediction

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

2026-05-29 · Zhenhao Yang, Xiaoshi Wu, Zhengyao Lv, Xiaoyu Shi 외 arxiv

Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horizon reasoning remains a key challenge. In…

Video Generation

Streaming Video Generation with Streaming Force Control

2026-06-05 · Hanhui Wang, Yiming Xie, Haiwen Feng, Zhaoyang Lv 외 arxiv

We introduce StreamForce, a streaming video generation framework that enables physically grounded control through continuous force inputs. Unlike prior video models that train separate models for different force types, a…

Video Generation

MotionClone: Training-Free Motion Cloning for Controllable Video Generation

2024-06-08 · Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong 외

Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning t…

DenoisingMotion GenerationMotion SynthesisText-to-Video Generation+1