paper-with-me

Papers

Evaluation of Text-to-Video Generation Models: A Dynamics Perspective

2024-07-01 · Mingxiang Liao, Hannan Lu, Xinyu Zhang, Fang Wan, Tianyu Wang, Yuzhong Zhao, WangMeng Zuo, Qixiang Ye, Jingdong Wang

Comprehensive and constructive evaluation protocols play an important role in the development of sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content continuity, yet largely ignore the dynamics of video content. Dynamics are an essential dimension for measuring the visual vividness and the honesty of video content to text prompts. In this study, we propose an effective evaluation protocol, termed DEVIL, which centers on the dynamics dimension to evaluate T2V models. For this purpose, we establish a new benchmark comprising text prompts that fully reflect multiple dynamics grades, and define a set of dynamics scores corresponding to various temporal granularities to comprehensively evaluate the dynamics of each generated video. Based on the new benchmark and the dynamics scores, we assess T2V models with the design of three metrics: dynamics range, dynamics controllability, and dynamics-based quality. Experiments show that DEVIL achieves a Pearson correlation exceeding 90% with human ratings, demonstrating its potential to advance T2V generation models. Code is available at https://github.com/MingXiangL/DEVIL.

📄 PDF Abstract BibTeX arXiv:2407.01094

Code (1)

mingxiangl/devil 공식 구현 pytorch

Tasks

Text-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective

2024-05-13 · Andrew Shin, Yusuke Mori, Kunitake Kaneko

Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models…

Text-to-Video GenerationVideo Generation

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation

2026-06-02 · Dingrui Wang, YuAn Wang, Jinkun Liu, Yue Zhang 외 arxiv

Recent world action models leverage video foundation models by aligning broad visual-dynamics priors with executable robot actions. We revisit this alignment from a distributional perspective. Existing formulations typic…

Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback

2024-12-03 · Hiroki Furuta, Heiga Zen, Dale Schuurmans, Aleksandra Faust 외

Large text-to-video models hold immense potential for a wide range of downstream applications. However, these models struggle to accurately depict dynamic object interactions, often resulting in unrealistic movements and…

ObjectOffline RLOptical Flow EstimationText-to-Video Generation+2

TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models

2025-11-26 · Jiaming He, Guanyu Hou, Hongwei Li, Zhicong Huang 외 arxiv

Text-to-Video (T2V) models are capable of synthesizing high-quality, temporally coherent dynamic video content, but the diverse generation also inherently introduces critical safety challenges. Existing safety evaluation…

Video GenerationText Generation

Beyond the Frame: Generating 360° Panoramic Videos from Perspective Videos

2025-04-10 · Rundong Luo, Matthew Wallingford, Ali Farhadi, Noah Snavely 외

360{\deg} videos have emerged as a promising medium to represent our dynamic visual world. Compared to the "tunnel vision" of standard cameras, their borderless field of view offers a more complete perspective of our sur…

Question AnsweringVideo GenerationVideo StabilizationVisual Question Answering