paper-with-me

Papers

Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

2024-12-08 · CVPR 2025 1 · Hyeonho Jeong, Chun-Hao Paul Huang, Jong Chul Ye, Niloy Mitra, Duygu Ceylan

While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across frames, breaking visual coherence. We hypothesize that this is because there is no explicit supervision in terms of spatial tracking at the feature level. We propose Track4Gen, a spatially aware video generator that combines video diffusion loss with point tracking across frames, providing enhanced spatial supervision on the diffusion features. Track4Gen merges the video generation and point tracking tasks into a single network by making minimal changes to existing video generation architectures. Using Stable Video Diffusion as a backbone, Track4Gen demonstrates that it is possible to unify video generation and point tracking, which are typically handled as separate tasks. Our extensive evaluations show that Track4Gen effectively reduces appearance drift, resulting in temporally stable and visually coherent video generation. Project page: hyeonho99.github.io/track4gen

📄 PDF Abstract BibTeX arXiv:2412.06016

Code (0)

등록된 구현이 없습니다.

Tasks

Point TrackingVideo Generation

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Point Prompting: Counterfactual Tracking with Video Diffusion Models

2025-10-13 · Ayush Shrivastava, Sanyam Mehta, Daniel Geng, Andrew Owens arxiv

Trackers and video generators solve closely related problems: the former analyze motion, while the latter synthesize it. We show that this connection enables pretrained video diffusion models to perform zero-shot point t…

Point Tracking

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

2026-05-12 · Jisu Nam, Jahyeok Koo, Soowon Son, Jaewoo Jung 외 arxiv

Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging a…

Scene Understanding3D Reconstruction

Teacher's Perception in the Classroom

2018-05-22 · Ömer Sümer, Patricia Goldberg, Kathleen Stürmer, Tina Seidel 외

The ability for a teacher to engage all students in active learning processes in classroom constitutes a crucial prerequisite for enhancing students achievement. Teachers' attentional processes provide important insights…

Active Learning

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

2026-07-02 · Long Vu, Tan Ngo, Animesh Karnewar, Amir Habibian 외 arxiv

Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approache…

Video Generation

TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models

2023-12-01 · Pengxiang Li, Kai Chen, Zhili Liu, Ruiyuan Gao 외

Despite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world mod…

Image ClassificationMulti-Object TrackingObjectobject-detection+3