paper-with-me

Papers

Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models

2024-02-24 · Chaoya Jiang, Hongrui Jia, Wei Ye, Mengfan Dong, Haiyang Xu, Ming Yan, Ji Zhang, Shikun Zhang

Large Vision Language Models exhibit remarkable capabilities but struggle with hallucinations inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms of objects, attributes, and relations but overlooked complex hallucinations that create an entire narrative around a fictional entity. In this paper, we introduce a refined taxonomy of hallucinations, featuring a new category: Event Hallucination. We then utilize advanced LLMs to generate and filter fine grained hallucinatory data consisting of various types of hallucinations, with a particular focus on event hallucinations, laying the groundwork for integrating discriminative and generative evaluation methods within our universal evaluation framework. The proposed benchmark distinctively assesses LVLMs ability to tackle a broad spectrum of hallucinations, making it a reliable and comprehensive tool for gauging LVLMs efficacy in handling hallucinations. We will release our code and data.

📄 PDF Abstract BibTeX arXiv:2402.15721

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationHallucination Evaluation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Mitigating Fine-Grained Hallucination by Fine-Tuning Large Vision-Language Models with Caption Rewrites

2023-12-04 · Lei Wang, Jiabang He, Shenshen Li, Ning Liu 외

Large language models (LLMs) have shown remarkable performance in natural language processing (NLP) tasks. To comprehend and execute diverse human instructions over image data, instruction-tuned large vision-language mod…

HallucinationHallucination EvaluationObjectObject Hallucination+1

Fine-grained Hallucination Detection and Editing for Language Models

2024-01-12 · Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang 외

Large language models (LMs) are prone to generate factual errors, which are often called hallucinations. In this paper, we introduce a comprehensive taxonomy of hallucinations and argue that hallucinations manifest in di…

HallucinationRetrieval

FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs

2026-03-20 · Zhihan Yin, Jianxin Liang, Yueqian Wang, Yifeng Yao 외 arxiv

Multimodal Large Language Models (MLLMs) suffer from hallucinations. Existing hallucination evaluation benchmarks are often limited by over-simplified tasks leading to saturated metrics, or insufficient diversity that fa…

AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs

2025-09-04 · Aisha Alansari, Hamzah Luqman arxiv

Recently, extensive research on the hallucination of the large language models (LLMs) has mainly focused on the English language. Despite the growing number of multilingual and Arabic-specific LLMs, evaluating LLMs' hall…

Generative Question Answering

SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs

2025-08-13 · Bei Yan, Zhiyuan Chen, Yuecong Min, Jie Zhang 외 arxiv

Despite rapid advances, Large Vision-Language Models (LVLMs) still suffer from hallucinations, i.e., generating content inconsistent with input or established world knowledge, which correspond to faithfulness and factual…