paper-with-me

Papers

Direct Motion Models for Assessing Generated Videos

2025-04-30 · Kelsey Allen, Carl Doersch, Guangyao Zhou, Mohammed Suhail, Danny Driess, Ignacio Rocco, Yulia Rubanova, Thomas Kipf, Mehdi S. M. Sajjadi, Kevin Murphy, Joao Carreira, Sjoerd van Steenkiste

A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -- an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond FVD by developing a metric which better measures plausible object interactions and motion. Our novel approach is based on auto-encoding point tracks and yields motion features that can be used to not only compare distributions of videos (as few as one generated and one ground truth, or as many as two datasets), but also for evaluating motion of single videos. We show that using point tracks instead of pixel reconstruction or action recognition features results in a metric which is markedly more sensitive to temporal distortions in synthetic data, and can predict human evaluations of temporal consistency and realism in generated videos obtained from open-source models better than a wide range of alternatives. We also show that by using a point track representation, we can spatiotemporally localize generative video inconsistencies, providing extra interpretability of generated video errors relative to prior work. An overview of the results and link to the code can be found on the project page: http://trajan-paper.github.io.

📄 PDF Abstract BibTeX arXiv:2505.00209

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

From Generated Human Videos to Physically Plausible Robot Trajectories

2025-12-04 · James Ni, Zekai Wang, Wei Lin, Amir Bar 외 arxiv

Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual robot control. To realize this potential,…

Zero-shot GeneralizationReinforcement LearningVideo Generation

What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality

2024-11-20 · Zihan Wang, Songlin Li, Lingyan Hao, Bowen Song 외

As video generation models advance rapidly, assessing the quality of generated videos has become increasingly critical. Existing metrics, such as Fr\'echet Video Distance (FVD), Inception Score (IS), and ClipSim, measure…

Video Generation

CamMimic: Zero-Shot Image To Camera Motion Personalized Video Generation Using Diffusion Models

2025-04-13 · Pooja Guhan, Divya Kothandaraman, Tsung-Wei Huang, Guan-Ming Su 외

We introduce CamMimic, an innovative algorithm tailored for dynamic video editing needs. It is designed to seamlessly transfer the camera motion observed in a given reference video onto any scene of the user's choice in …

Video EditingVideo Generation

SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models

2026-06-26 · Ruoyu Wang, Jialun Liu, Huayang Huang, Haibin Huang 외 arxiv

Recent advances in video diffusion models have greatly improved visual fidelity, yet their generated motions often violate physical plausibility. We observe a common kinematic failure, "motion entanglement", the unintend…

Temporal Realism Evaluation of Generated Videos Using Compressed-Domain Motion Vectors

2025-11-17 · Mert Onur Cakiroglu, Idil Bilge Altun, Zhihe Lu, Mehmet Dalkilic 외 arxiv

Temporal realism remains a central weakness of current generative video models, as most evaluation metrics prioritize spatial appearance and offer limited sensitivity to motion. We introduce a scalable, model-agnostic fr…