Papers Hallucination Evaluation
“Hallucination Evaluation” 태그가 달린 논문 49편 · 필터 해제
HalluSegBench: Counterfactual Visual Reasoning for Segmentation Hallucination Evaluation
Recent progress in vision-language segmentation has significantly advanced grounded visual understanding. However, these models often exhibit hallucinations by producing segmentation masks for objects not grounded in the…
counterfactualCounterfactual ReasoningHallucinationHallucination Evaluation+4KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
Large Language Models (LLMs), particularly slow-thinking models, often exhibit severe hallucination, outputting incorrect content due to an inability to accurately recognize knowledge boundaries during reasoning. While R…
HallucinationHallucination Evaluationreinforcement-learningReinforcement Learning+1MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
Large Language Models (LLMs) have inherent limitations of faithfulness and factuality, commonly referred to as hallucinations. Several benchmarks have been developed that provide a test bed for factuality evaluation with…
Fact CheckingHallucinationHallucination EvaluationKnowledge Graphs+5Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
Hallucinations remain a persistent challenge for LLMs. RAG aims to reduce hallucinations by grounding responses in contexts. However, even when provided context, LLMs still frequently introduce unsupported information or…
BenchmarkingHallucinationHallucination EvaluationRAGMitigating Image Captioning Hallucinations in Vision-Language Models
Hallucinations in vision-language models (VLMs) hinder reliability and real-world applicability, usually stemming from distribution shifts between pretraining data and test samples. Existing solutions, such as retraining…
HallucinationHallucination EvaluationImage CaptioningLanguage Modeling+2Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due t…
HallucinationHallucination EvaluationVisual Question Answering (VQA)Visual ReasoningReal-Time Evaluation Models for RAG: Who Detects Hallucinations Best?
This article surveys Evaluation models to automatically detect hallucinations in Retrieval-Augmented Generation (RAG), and presents a comprehensive benchmark of their performance across six RAG applications. Methods incl…
HallucinationHallucination EvaluationLanguage ModelingLanguage Modelling+3Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
Preference alignment has emerged as an effective strategy to enhance the performance of Multimodal Large Language Models (MLLMs) following supervised fine-tuning. While existing preference alignment methods predominantly…
HallucinationHallucination EvaluationQuestion AnsweringVisual Question AnsweringExploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
The hallucination of large multimodal models (LMMs), providing responses that appear correct but are actually incorrect, limits their reliability and applicability. This paper aims to study the hallucination problem of L…
HallucinationHallucination EvaluationVideo UnderstandingEvaluating LLMs' Assessment of Mixed-Context Hallucination Through the Lens of Summarization
With the rapid development of large language models (LLMs), LLM-as-a-judge has emerged as a widely adopted approach for text quality evaluation, including hallucination evaluation. While previous studies have focused exc…
HallucinationHallucination EvaluationTreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
Large language models (LLMs) now achieve near-human performance on standard math word problem benchmarks (e.g., GSM8K), yet their true reasoning ability remains disputed. A key concern is that models often produce confid…
Dataset GenerationGSM8KHallucinationHallucination Evaluation+1Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
Multimodal Large Language Models (MLLMs) are known to hallucinate, which limits their practical applications. Recent works have attempted to apply Direct Preference Optimization (DPO) to enhance the performance of MLLMs,…
HallucinationHallucination EvaluationLanguage ModelingLanguage Modelling+2DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
We introduce DAHL, a benchmark dataset and automated evaluation system designed to assess hallucination in long-form text generation, specifically within the biomedical domain. Our benchmark dataset, meticulously curated…
FormHallucinationHallucination EvaluationMultiple-choice+1DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark
With the rapid development of large vision language models (LVLMs), these models have shown excellent results in various multimodal tasks. Since LVLMs are prone to hallucinations and there are currently few datasets and …
Data AugmentationHallucinationHallucination EvaluationUnified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
Despite the outstanding performance in vision-language reasoning, Large Vision-Language Models (LVLMs) might generate hallucinated contents that do not exist in the given image. Most existing LVLM hallucination benchmark…
HallucinationHallucination EvaluationObjectObject Hallucination+2A Survey of Hallucination in Large Visual Language Models
The Large Visual Language Models (LVLMs) enhances user interaction and enriches user experience by integrating visual modality on the basis of the Large Language Models (LLMs). It has demonstrated their powerful informat…
HallucinationHallucination EvaluationSurveyLongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
Hallucination, a phenomenon where multimodal large language models~(MLLMs) tend to generate textual responses that are plausible but unaligned with the image, has become one major hurdle in various MLLM-related applicati…
HallucinationHallucination EvaluationMultiple-choiceTLDR: Token-Level Detective Reward Model for Large Vision Language Models
Although reward models have been successful in improving multimodal large language models, the reward models themselves remain brutal and contain minimal information. Notably, existing reward models only mimic human anno…
HallucinationHallucination EvaluationEffectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
Recent studies have shown that Vision Language Large Models (VLLMs) may output content not relevant to the input images. This problem, called the hallucination phenomenon, undoubtedly degrades VLLM performance. Therefore…
HallucinationHallucination EvaluationImage CaptioningResponse GenerationFIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs
The rapid development of Large Vision-Language Models (LVLMs) often comes with widespread hallucination issues, making cost-effective and comprehensive assessments increasingly vital. Current approaches mainly rely on co…
HallucinationHallucination Evaluation