paper-with-me

Papers

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs

2024-11-25 · Wey Yeh Choong, Yangyang Guo, Mohan Kankanhalli

Vision Large Language Models (VLLMs) are widely acknowledged to be prone to hallucination. Existing research addressing this problem has primarily been confined to image inputs, with limited exploration of video-based hallucinations. Furthermore, current evaluation methods fail to capture nuanced errors in generated responses, which are often exacerbated by the rich spatiotemporal dynamics of videos. To address this, we introduce VidHal, a benchmark specially designed to evaluate video-based hallucinations in VLLMs. VidHal is constructed by bootstrapping video instances across common temporal aspects. A defining feature of our benchmark lies in the careful creation of captions which represent varying levels of hallucination associated with each video. To enable fine-grained evaluation, we propose a novel caption ordering task requiring VLLMs to rank captions by hallucinatory extent. We conduct extensive experiments on VidHal and comprehensively evaluate a broad selection of models. Our results uncover significant limitations in existing VLLMs regarding hallucination generation. Through our benchmark, we aim to inspire further research on 1) holistic understanding of VLLM capabilities, particularly regarding hallucination, and 2) extensive development of advanced VLLMs to alleviate this problem.

📄 PDF Abstract BibTeX arXiv:2411.16771

Code (1)

lookuz/vidhal 공식 구현 pytorch

Tasks

BenchmarkingHallucination

Similar Papers 제목 키워드 기반

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

2024-12-04 · CVPR 2025 1 · Chaoyu Li, Eun Woo Im, Pooyan Fazli

Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. However, hallucination, where models generate …

HallucinationInstruction FollowingSemantic SimilaritySemantic Textual Similarity+1

Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators

2026-09-09 · Xinyu Chen, Adnan Mahmood, Mark Dras arxiv

Video-language models and video agents can produce hallucinations that conflict with spatiotemporal evidence. Existing benchmarks mainly evaluate model hallucinations, and heterogeneous mechanisms make detector reliabili…

Video Question AnsweringVideo Captioning

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

2026-02-19 · Zixu Cheng, Da Li, Jian Hu, Yuhang Zang 외 arxiv

Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language Models (MLLMs) are prone to severe temp…

Visual Grounding

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

2026-05-31 · Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu, Tae-Hyun Oh arxiv

Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate…

HypoTermQA: Hypothetical Terms Dataset for Benchmarking Hallucination Tendency of LLMs

2024-02-25 · Cem Uluoglakci, Tugba Taskaya Temizel

Hallucinations pose a significant challenge to the reliability and alignment of Large Language Models (LLMs), limiting their widespread acceptance beyond chatbot applications. Despite ongoing efforts, hallucinations rema…

BenchmarkingChatbotHallucinationLanguage Modeling+1