paper-with-me

홈 › Papers

Trajectory Attention for Fine-grained Video Motion Control

2024-11-28 · Zeqi Xiao, Wenqi Ouyang, Yifan Zhou, Shuai Yang, Lei Yang, Jianlou Si, Xingang Pan

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a novel approach that performs attention along available pixel trajectories for fine-grained camera motion control. Unlike existing methods that often yield imprecise outputs or neglect temporal correlations, our approach possesses a stronger inductive bias that seamlessly injects trajectory information into the video generation process. Importantly, our approach models trajectory attention as an auxiliary branch alongside traditional temporal attention. This design enables the original temporal attention and the trajectory attention to work in synergy, ensuring both precise motion control and new content generation capability, which is critical when the trajectory is only partially available. Experiments on camera motion control for images and videos demonstrate significant improvements in precision and long-range consistency while maintaining high-quality generation. Furthermore, we show that our approach can be extended to other video motion control tasks, such as first-frame-guided video editing, where it excels in maintaining content consistency over large spatial and temporal ranges.

📄 PDF Abstract BibTeX arXiv:2411.19324

Code (0)

등록된 구현이 없습니다.

Tasks

Inductive BiasVideo EditingVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

2025-10-01 · Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li 외 arxiv

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Visi…

Trajectory ForecastingTrajectory PredictionVideo Generation

Interpretable Video Captioning via Trajectory Structured Localization

2018-06-01 · CVPR 2018 6 · Xian Wu, Guanbin Li, Qingxing Cao, Qingge Ji 외

Automatically describing open-domain videos with natural language are attracting increasing interest in the field of artificial intelligence. Most existing methods simply borrow ideas from image captioning and obtain a c…

DecoderImage CaptioningSentenceVideo Captioning+1

Track and Caption Any Motion: Query-Free Motion Discovery and Description in Videos

2025-12-11 · Bishoy Galoaa, Sarah Ostadabbas arxiv

We propose Track and Caption Any Motion (TCAM), a motion-centric framework for automatic video understanding that discovers and describes motion patterns without user queries. Understanding videos in challenging conditio…

Text Retrieval

MotionPro: A Precise Motion Controller for Image-to-Video Generation

2025-05-26 · CVPR 2025 1 · Zhongwei Zhang, Fuchen Long, Zhaofan Qiu, Yingwei Pan 외

Animating images with interactive motion control has garnered popularity for image-to-video (I2V) generation. Modern approaches typically rely on large Gaussian kernels to extend motion trajectories as condition without …

DenoisingImage to Video GenerationMotion SynthesisVideo Denoising+1

Follow Your Track: Precise Skeleton Animation Controlled by 3D Trajectories

2026-06-24 · Yueting Liu, Yanqin Jiang, Nian Liu, Jingmen Zhou 외 arxiv

4D generation aims to animate 3D objects with realistic motion, holding great promise for applications. Existing methods typically decouple 3D asset generation from motion synthesis: acquire a 3D asset, prepare a structu…

Motion Synthesis