paper-with-me

홈 › Papers

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

2026-07-15 · Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang hf

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.

📄 PDF Abstract BibTeX arXiv:2607.14202

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

KEMP: Keyframe-Based Hierarchical End-to-End Deep Model for Long-Term Trajectory Prediction

2022-05-10 · QIUJING LU, Weiqiao Han, Jeffrey Ling, Minfa Wang 외

Predicting future trajectories of road agents is a critical task for autonomous driving. Recent goal-based trajectory prediction methods, such as DenseTNT and PECNet, have shown good performance on prediction tasks on pu…

Autonomous DrivingPredictionTrajectory Prediction

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

2025-01-01 · CVPR 2025 1 · Antoni Bigata, Michał Stypułkowski, Rodrigo Mira, Stella Bounareli 외

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate …

Anchoring and Rescaling Attention for Semantically Coherent Inbetweening

2026-03-18 · Tae Eun Choi, Sumin Shim, Junhyeok Kim, Seong Jae Hwang arxiv

Generative inbetweening (GI) seeks to synthesize realistic intermediate frames between the first and last keyframes beyond mere interpolation. As sequences become sparser and motions larger, previous GI models struggle w…

SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation

2026-03-17 · Jiongze Yu, Xiangbo Gao, Pooja Verlani, Akshay Gadde 외 arxiv

Video Super-Resolution (VSR) aims to restore high-quality video frames from low-resolution (LR) estimates, yet most existing VSR approaches behave like black boxes at inference time: users cannot reliably correct unexpec…

Image Super-ResolutionVideo Super-ResolutionStyle Transfer

Feed-forward Motion In-betweening for Any 4D

2026-06-20 · Hiroki Nishizawa, Hubert P. H. Shum, Yoshihiro Fukuhara, Hirokatsu Kataoka 외 arxiv

4D dynamics (3D geometry evolving over time) is a fundamental representation of the physical world and plays a crucial role in world modeling (e.g., animation and games). Owing to the scarcity of large-scale, long-horizo…