paper-with-me

홈 › Papers

Towards A Better Metric for Text-to-Video Generation

2024-01-15 · Jay Zhangjie Wu, Guian Fang, HaoNing Wu, Xintao Wang, Yixiao Ge, Xiaodong Cun, David Junhao Zhang, Jia-Wei Liu, YuChao Gu, Rui Zhao, Weisi Lin, Wynne Hsu, Ying Shan, Mike Zheng Shou

Generative models have demonstrated remarkable capability in synthesizing high-quality text, images, and videos. For video generation, contemporary text-to-video models exhibit impressive capabilities, crafting visually stunning videos. Nonetheless, evaluating such videos poses significant challenges. Current research predominantly employs automated metrics such as FVD, IS, and CLIP Score. However, these metrics provide an incomplete analysis, particularly in the temporal assessment of video content, thus rendering them unreliable indicators of true video quality. Furthermore, while user studies have the potential to reflect human perception accurately, they are hampered by their time-intensive and laborious nature, with outcomes that are often tainted by subjective bias. In this paper, we investigate the limitations inherent in existing metrics and introduce a novel evaluation pipeline, the Text-to-Video Score (T2VScore). This metric integrates two pivotal criteria: (1) Text-Video Alignment, which scrutinizes the fidelity of the video in representing the given text description, and (2) Video Quality, which evaluates the video's overall production caliber with a mixture of experts. Moreover, to evaluate the proposed metrics and facilitate future improvements on them, we present the TVGE dataset, collecting human judgements of 2,543 text-to-video generated videos on the two criteria. Experiments on the TVGE dataset demonstrate the superiority of the proposed T2VScore on offering a better metric for text-to-video generation.

📄 PDF Abstract BibTeX arXiv:2401.07781

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-ExpertsText-to-Video GenerationVideo AlignmentVideo Generation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

2024-07-19 · CVPR 2025 1 · Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu 외

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also …

AttributeLanguage ModelingLanguage ModellingLarge Language Model+3

Knowledge-Intensive Video Generation

2026-05-31 · Chenxu Wang, Mingda Chen arxiv

Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness. We introduce knowledge-intensive video generation (KIVI), where models generate videos…

Text-to-Video Generation

Exploring the Distributed Video Coding in a Quality Assessment Context

2018-03-13

In the popular video coding trend, the encoder has the task to exploit both spatial and temporal redundancies present in the video sequence, which is a complex procedure. As a result almost all video encoders have five t…

DecoderMotion CompensationMotion EstimationVideo Compression

StoryBench: A Multifaceted Benchmark for Continuous Story Visualization

2023-08-22 · NeurIPS 2023 11 · Emanuele Bugliarello, Hernan Moraldo, Ruben Villegas, Mohammad Babaeizadeh 외

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Cr…

Story ContinuationStory GenerationStory VisualizationVideo Generation

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

2023-09-28 · Guy Yariv, Itai Gat, Sagie Benaim, Lior Wolf 외

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally w…

Text-to-Video GenerationVideo Generation