paper-with-me

홈 › Papers

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

2026-06-23 · Linpeng Huang, Weixing Chen, Zexin Chen, Yang Liu, Liang Lin arxiv

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctness, while the grounding of predictions in relevant video evidence remains largely unexamined. This disconnect between answer generation and evidence understanding motivates the construction of the Evidence-Grounded Video Question Answering Benchmark (EG-VQA), an open-ended evaluation protocol in which each QA pair is explicitly annotated with supporting temporal evidence, thereby requiring joint reasoning and precise evidence localization. EG-VQA is comprised of 2,067 videos and 11,838 QA pairs with fine-grained evidence annotations. To evaluate predicted evidence, Evidence-Grounded F1 (EG-F1) is introduced as a unified metric in which temporal alignment and semantic consistency against ground-truth evidence are jointly measured. Experimental evaluation reveals that even strong proprietary models struggle to accurately ground their predictions, exposing a fundamental discrepancy between answer correctness and faithful evidence localization. To bridge this gap, EG-Reasoner, an evidence-grounded reasoning model trained with explicit supervision, is proposed. State-of-the-art performance is achieved among open-source models, with results competitive against proprietary systems, particularly pronounced gains are observed on reasoning-intensive tasks such as counterfactual questions. These findings demonstrate that scaling alone is insufficient for robust video understanding and that structured evidence supervision is essential for the development of more reliable and interpretable VideoQA systems.

📄 PDF Abstract BibTeX arXiv:2606.24797

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question AnsweringAnswer Generation

Similar Papers 제목 키워드 기반

Perception Test 2023: A Summary of the First Challenge And Outcome

2023-12-20 · Joseph Heyward, João Carreira, Dima Damen, Andrew Zisserman 외

The First Perception Test challenge was held as a half-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2023, with the goal of benchmarking state-of-the-art video models on the recen…

BenchmarkingGrounded Video Question AnsweringMultiple-choiceObject Tracking+3

Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning

2025-01-09 · CVPR 2025 1 · Huabin Liu, Filip Ilievski, Cees G. M. Snoek

This paper proposes the first video-grounded entailment tree reasoning method for commonsense video question answering (VQA). Despite the remarkable progress of large visual-language models (VLMs), there are growing conc…

BenchmarkingQuestion AnsweringVideo Question AnsweringVisual Question Answering (VQA)

Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark

2024-11-29 · Joseph Heyward, João Carreira, Dima Damen, Andrew Zisserman 외

Following the successful 2023 edition, we organised the Second Perception Test challenge as a half-day workshop alongside the IEEE/CVF European Conference on Computer Vision (ECCV) 2024, with the goal of benchmarking sta…

BenchmarkingGrounded Video Question AnsweringMultiple-choiceObject Tracking+5

Evidence-Backed Video Question Answering

2026-07-13 · Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu 외 hf

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on…

Video Question AnsweringObject SegmentationVisual Grounding

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

2026-05-22 · Mingfang Zhang, Jingjing Pan, Ashutosh Kumar, Rajat Saini 외 arxiv

Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing ben…

Video Question Answering