paper-with-me

홈 › Papers

Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising

2025-11-09 · Assaf Singer, Noam Rotstein, Amir Mann, Ron Kimmel, Or Litany arxiv

Diffusion-based video generation can create realistic videos, yet existing image- and text-based conditioning fails to offer precise motion control. Prior methods for motion-conditioned synthesis typically require model-specific fine-tuning, which is computationally expensive and restrictive. We introduce Time-to-Move (TTM), a training-free, plug-and-play framework for motion- and appearance-controlled video generation with image-to-video (I2V) diffusion models. Our key insight is to use crude reference animations obtained through user-friendly manipulations such as cut-and-drag or depth-based reprojection. Motivated by SDEdit's use of coarse layout cues for image editing, we treat the crude animations as coarse motion cues and adapt the mechanism to the video domain. We preserve appearance with image conditioning and introduce dual-clock denoising, a region-dependent strategy that enforces strong alignment in motion-specified regions while allowing flexibility elsewhere, balancing fidelity to user intent with natural dynamics. This lightweight modification of the sampling process incurs no additional training or runtime cost and is compatible with any backbone. Extensive experiments on object and camera motion benchmarks show that TTM matches or exceeds existing training-based baselines in realism and motion control. Beyond this, TTM introduces a unique capability: precise appearance control through pixel-level conditioning, exceeding the limits of text-only prompting. Visit our project page for video examples and code: https://time-to-move.github.io/.

📄 PDF Abstract BibTeX arXiv:2511.08633

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationImage Editing

Similar Papers 제목 키워드 기반

EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras

2016-09-23 · Helge Rhodin, Christian Richardt, Dan Casas, Eldar Insafutdinov 외

Marker-based and marker-less optical skeletal motion-capture methods use an outside-in arrangement of cameras placed around a scene, with viewpoints converging on the center. They often create discomfort by possibly need…

Pose EstimationVocal Bursts Valence Prediction

CALM: Conditional Adversarial Latent Models for Directable Virtual Characters

2023-05-02 · Chen Tessler, Yoni Kasten, Yunrong Guo, Shie Mannor 외

In this work, we present Conditional Adversarial Latent Models (CALM), an approach for generating diverse and directable behaviors for user-controlled interactive virtual characters. Using imitation learning, CALM learns…

DiversityImitation Learning

Dimensionality Reduction and Motion Clustering during Activities of Daily Living: 3, 4, and 7 Degree-of-Freedom Arm Movements

2020-02-17 · Yuri Gloumakov, Adam J. Spiers, Aaron M. Dollar

The wide variety of motions performed by the human arm during daily tasks makes it desirable to find representative subsets to reduce the dimensionality of these movements for a variety of applications, including the des…

ClusteringDimensionality ReductionDynamic Time Warping

Versatile Multimodal Controls for Expressive Talking Human Animation

2025-03-10 · Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li 외

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not onl…

Human Animation

Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video

2026-05-14 · Yifan Wang, Tong He arxiv

Camera-controlled video generation has made substantial progress, enabling generated videos to follow prescribed viewpoint trajectories. However, existing methods usually learn camera-specific conditioning through camera…

Video Generation