paper-with-me

홈 › Papers

Spotlight: Identifying and Localizing Video Generation Errors Using VLMs

2025-11-22 · Aditya Chinchure, Sahithya Ravi, Pushkar Shukla, Vered Shwartz, Leonid Sigal arxiv

Current text-to-video models (T2V) can generate high-quality, temporally coherent, and visually realistic videos. Nonetheless, errors still often occur, and are more nuanced and local compared to the previous generation of T2V models. While current evaluation paradigms assess video models across diverse dimensions, they typically evaluate videos holistically without identifying when specific errors occur or describing their nature. We address this gap by introducing Spotlight, a novel task aimed at localizing and explaining video-generation errors. We generate 600 videos using 200 diverse textual prompts and three state-of-the-art video generators (Veo 3, Seedance, and LTX-2), and annotate over 1600 fine-grained errors across six types, including motion, physics, and prompt adherence. We observe that adherence and physics errors are predominant and persist across longer segments, whereas appearance-disappearance and body pose errors manifest in shorter segments. We then evaluate current VLMs on Spotlight and find that VLMs lag significantly behind humans in error identification and localization in videos. We propose inference-time strategies to probe the limits of current VLMs on our task, improving performance by nearly 2x. Our task paves a way forward to building fine-grained evaluation tools and more sophisticated reward models for video generators.

📄 PDF Abstract BibTeX arXiv:2511.18102

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

The Spotlight: A General Method for Discovering Systematic Errors in Deep Learning Models

2021-07-01 · Greg d'Eon, Jason d'Eon, James R. Wright, Kevin Leyton-Brown

Supervised learning models often make systematic errors on rare subsets of the data. When these subsets correspond to explicit labels in the data (e.g., gender, race) such poor performance can be identified straightforwa…

Recommendation Systems

Active Light Modulation to Counter Manipulation of Speech Visual Content

2025-04-30 · Hadleigh Schwartz, Xiaofeng Yan, Charles J. Carver, Xia Zhou

High-profile speech videos are prime targets for falsification, owing to their accessibility and influence. This work proposes Spotlight, a low-overhead and unobtrusive system for protecting live speech videos from visua…

The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective

2024-05-13 · Andrew Shin, Yusuke Mori, Kunitake Kaneko

Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models…

Text-to-Video GenerationVideo Generation

UAL-Bench: The First Comprehensive Unusual Activity Localization Benchmark

2024-10-02 · Hasnat Md Abdullah, Tian Liu, Kangda Wei, Shu Kong 외

Localizing unusual activities, such as human errors or surveillance incidents, in videos holds practical significance. However, current video understanding models struggle with localizing these unusual events likely beca…

Unusual Activity LocalizationVideo Understanding

VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs

2024-06-14 · Rohit Bharadwaj, Hanan Gani, Muzammal Naseer, Fahad Shahbaz Khan 외

The recent developments in Large Multi-modal Video Models (Video-LMMs) have significantly enhanced our ability to interpret and analyze video data. Despite their impressive capabilities, current Video-LMMs have not been …

Anomaly DetectionBenchmarkingQuestion AnsweringText-to-Video Generation+2