paper-with-me

Papers

No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs

2026-06-30 · Haojian Huang, Harold Haodong Chen, Meng Luo, Junjia Du, Shanqing Xu, Ziheng Chen, Yanxiang Huang, Yinchuan Li, Ying-Cong Chen arxiv

We introduce VidPair-Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. Unlike previous benchmarks that primarily rely on text-based perturbations or adversarial questions while neglecting the consistency of visual backgrounds, VidPair-Halluc features video pairs with highly similar backgrounds but distinctly different foreground semantics, enabling precise attribution of model errors to genuine hallucination rather than background variation. The benchmark is constructed through PairFlow, a pipeline that leverages recent advances in text-to-image and video generation to systematically compose stories, generate coherent video clips, and assemble them into adversarial pairs. Covering both spatial and temporal reasoning across ten semantic aspects, VidPair-Halluc comprises 1K high-quality adversarial video pairs and 11K spatio-temporal QA pairs with control over background and foreground variations. Evaluations on mainstream LVMs show persistent difficulty with robust fine-grained video understanding in adversarial settings, and code and data are available at the https://jethrojames.github.io/VidPair-Halluc/.

📄 PDF Abstract BibTeX arXiv:2606.31933

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models

2026-03-16 · Amira Guesmi, Muhammad Shafique arxiv

Vision-language models (VLMs) have recently shown remarkable capabilities in visual understanding and generation, but remain vulnerable to adversarial manipulations of visual content. Prior object-hiding attacks primaril…

When to Lock Attention: Training-Free KV Control in Video Diffusion

2026-03-10 · Tianyi Zeng, Jincheng Gao, Tianyi Wang, Zijie Meng 외 arxiv

Maintaining background consistency while enhancing foreground quality remains a core challenge in video editing. Injecting full-image information often leads to background artifacts, whereas rigid background locking seve…

Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

2026-08-02 · Timothee Mickus, Claudio Savelli, Eduardo Calò, Emilio Raimond 외 arxiv

In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make b…

A Camera-Native Talking-Head Video Dataset for Various Computer Vision Tasks

2026-03-23 · Babak Naderi, Ross Cutler, Nabakumar Singh Khongbantabam arxiv

Talking-head videos constitute a predominant content type in real-time communication, yet publicly available datasets for video processing research in this domain remain scarce and limited in signal fidelity. In this pap…

Exploring Audio Hallucination in Egocentric Video Understanding

2026-04-26 · Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja 외 arxiv

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstable or occluded due to continuous camera …