paper-with-me

홈 › Papers

INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs

2026-03-12 · Junqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan, Xilin Chen arxiv

Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or verifiable world knowledge (factuality). Existing benchmarks provide limited coverage of factuality hallucinations and predominantly evaluate models only in clean settings. We introduce \textsc{INFACT}, a diagnostic benchmark comprising 9{,}800 QA instances with fine-grained taxonomies for faithfulness and factuality, spanning real and synthetic videos. \textsc{INFACT} evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items. Reliability under induced modes is quantified using Resist Rate (RR) and Temporal Sensitivity Score (TSS). Experiments on 14 representative Video-LLMs reveal that higher Base-mode accuracy does not reliably translate to higher reliability in the induced modes, with evidence corruption reducing stability and temporal intervention yielding the largest degradation. Notably, many open-source baselines exhibit near-zero TSS on factuality, indicating pronounced temporal inertia on order-sensitive questions.

📄 PDF Abstract BibTeX arXiv:2603.11481

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models

2024-03-01 · Kedi Chen, Qin Chen, Jie zhou, Yishen He 외

Since large language models (LLMs) achieve significant success in recent years, the hallucination issue remains a challenge, numerous benchmarks are proposed to detect the hallucination. Nevertheless, some of these bench…

HallucinationHallucination EvaluationSentence

SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs

2025-08-13 · Bei Yan, Zhiyuan Chen, Yuecong Min, Jie Zhang 외 arxiv

Despite rapid advances, Large Vision-Language Models (LVLMs) still suffer from hallucinations, i.e., generating content inconsistent with input or established world knowledge, which correspond to faithfulness and factual…

PlainQAFact: Automatic Factuality Evaluation Metric for Biomedical Plain Language Summaries Generation

2025-03-11 · Zhiwen You, Yue Guo

Hallucinated outputs from language models pose risks in the medical domain, especially for lay audiences making health-related decisions. Existing factuality evaluation methods, such as entailment- and question-answering…

Question Answering

Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness

2024-03-30 · Baolong Bi, Shenghua Liu, Yiwei Wang, Lingrui Mei 외

As the modern tools of choice for text understanding and generation, large language models (LLMs) are expected to accurately output answers by leveraging the input context. This requires LLMs to possess both context-fait…

knowledge editing

HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

2024-06-11 · Wen Luo, Tianshu Shen, Wei Li, Guangyue Peng 외

Large Language Models (LLMs) have significantly advanced the field of Natural Language Processing (NLP), achieving remarkable performance across diverse tasks and enabling widespread real-world applications. However, LLM…

HallucinationHallucination EvaluationLanguage ModellingSentence