paper-with-me

홈 › Papers

VideoLLM Benchmarks and Evaluation: A Survey

2025-05-03 · Yogesh Kumar

The rapid development of Large Language Models (LLMs) has catalyzed significant advancements in video understanding technologies. This survey provides a comprehensive analysis of benchmarks and evaluation methodologies specifically designed or used for Video Large Language Models (VideoLLMs). We examine the current landscape of video understanding benchmarks, discussing their characteristics, evaluation protocols, and limitations. The paper analyzes various evaluation methodologies, including closed-set, open-set, and specialized evaluations for temporal and spatiotemporal understanding tasks. We highlight the performance trends of state-of-the-art VideoLLMs across these benchmarks and identify key challenges in current evaluation frameworks. Additionally, we propose future research directions to enhance benchmark design, evaluation metrics, and protocols, including the need for more diverse, multimodal, and interpretability-focused benchmarks. This survey aims to equip researchers with a structured understanding of how to effectively evaluate VideoLLMs and identify promising avenues for advancing the field of video understanding with large language models.

📄 PDF Abstract BibTeX arXiv:2505.03829

Code (0)

등록된 구현이 없습니다.

Tasks

SurveyVideo Understanding

Similar Papers 제목 키워드 기반

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

2026-09-09 · Killian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi arxiv

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their …

Question Answering

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning

2025-01-12 · Ji Soo Lee, Jongha Kim, Jeehye Na, Jinyoung Park 외

Despite the advancements of Video Large Language Models (VideoLLMs) in various tasks, they struggle with fine-grained temporal understanding, such as Dense Video Captioning (DVC). DVC is a complicated task of describing …

Dense Video CaptioningVideo CaptioningVideo GroundingVideo Segmentation+2

VideoSTF: Stress-Testing Output Repetition in Video Large Language Models

2026-02-11 · Yuxin Cao, Wei Song, Shangzhi Xu, Jingling Xue 외 arxiv

Video Large Language Models (VideoLLMs) have recently achieved strong performance in video understanding tasks. However, we identify a previously underexplored generation failure: severe output repetition, where models d…

ARGUS: Hallucination and Omission Evaluation in Video-LLMs

2025-06-09 · Ruchit Rawal, Reza Shirkavand, Heng Huang, Gowthami Somepalli 외

Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple-choice questions. Unfortunately, VideoLLMs hallucinat…

DescriptiveFormHallucinationMultiple-choice+2

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models

2026-07-01 · Jiale Li, Sihan Chen, Mengyuan Liu arxiv

Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsistent with visual evidence. Existing benchmarks mainly focus on object …

Action Recognition