paper-with-me

Papers Hallucination Evaluation

“Hallucination Evaluation” 태그가 달린 논문 49편 · 필터 해제

HalluSegBench: Counterfactual Visual Reasoning for Segmentation Hallucination Evaluation

2025-06-26 · Xinzhuo Li, Adheesh Juvekar, Xingyou Liu, Muntasir Wahed 외

Recent progress in vision-language segmentation has significantly advanced grounded visual understanding. However, these models often exhibit hallucinations by producing segmentation masks for objects not grounded in the…

counterfactualCounterfactual ReasoningHallucinationHallucination Evaluation+4

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality

2025-06-24 · Baochang Ren, Shuofei Qiao, Wenhao Yu, Huajun Chen 외

Large Language Models (LLMs), particularly slow-thinking models, often exhibit severe hallucination, outputting incorrect content due to an inability to accurately recognize knowledge boundaries during reasoning. While R…

HallucinationHallucination Evaluationreinforcement-learningReinforcement Learning+1

MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations

2025-05-20 · Ernests Lavrinovics, Russa Biswas, Katja Hose, Johannes Bjerva

Large Language Models (LLMs) have inherent limitations of faithfulness and factuality, commonly referred to as hallucinations. Several benchmarks have been developed that provide a test bed for factuality evaluation with…

Fact CheckingHallucinationHallucination EvaluationKnowledge Graphs+5

Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards

2025-05-07 · Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo 외

Hallucinations remain a persistent challenge for LLMs. RAG aims to reduce hallucinations by grounding responses in contexts. However, even when provided context, LLMs still frequently introduce unsupported information or…

BenchmarkingHallucinationHallucination EvaluationRAG

Mitigating Image Captioning Hallucinations in Vision-Language Models

2025-05-06 · Fei Zhao, Chengcui Zhang, Runlin Zhang, Tianyang Wang 외

Hallucinations in vision-language models (VLMs) hinder reliability and real-world applicability, usually stemming from distribution shifts between pretraining data and test samples. Existing solutions, such as retraining…

HallucinationHallucination EvaluationImage CaptioningLanguage Modeling+2

Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs

2025-04-30 · Dung Nguyen, Minh Khoi Ho, Huy Ta, Thanh Tam Nguyen 외

Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due t…

HallucinationHallucination EvaluationVisual Question Answering (VQA)Visual Reasoning

Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best?

2025-03-27 · Ashish Sardana

This article surveys Evaluation models to automatically detect hallucinations in Retrieval-Augmented Generation (RAG), and presents a comprehensive benchmark of their performance across six RAG applications. Methods incl…

HallucinationHallucination EvaluationLanguage ModelingLanguage Modelling+3

Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs

2025-03-26 · Zitian Wang, Yue Liao, Kang Rong, Fengyun Rao 외

Preference alignment has emerged as an effective strategy to enhance the performance of Multimodal Large Language Models (MLLMs) following supervised fine-tuning. While existing preference alignment methods predominantly…

HallucinationHallucination EvaluationQuestion AnsweringVisual Question Answering

Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation

2025-03-25 · Hongcheng Gao, Jiashu Qu, Jingyi Tang, Baolong Bi 외

The hallucination of large multimodal models (LMMs), providing responses that appear correct but are actually incorrect, limits their reliability and applicability. This paper aims to study the hallucination problem of L…

HallucinationHallucination EvaluationVideo Understanding

Evaluating LLMs' Assessment of Mixed-Context Hallucination Through the Lens of Summarization

2025-03-03 · Siya Qi, Rui Cao, Yulan He, Zheng Yuan

With the rapid development of large language models (LLMs), LLM-as-a-judge has emerged as a widely adopted approach for text quality evaluation, including hallucination evaluation. While previous studies have focused exc…

HallucinationHallucination Evaluation

TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation

2025-02-19 · Jialin Ouyang

Large language models (LLMs) now achieve near-human performance on standard math word problem benchmarks (e.g., GSM8K), yet their true reasoning ability remains disputed. A key concern is that models often produce confid…

Dataset GenerationGSM8KHallucinationHallucination Evaluation+1

Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization

2024-11-15 · Yuhan Fu, Ruobing Xie, Xingwu Sun, Zhanhui Kang 외

Multimodal Large Language Models (MLLMs) are known to hallucinate, which limits their practical applications. Recent works have attempted to apply Direct Preference Optimization (DPO) to enhance the performance of MLLMs,…

HallucinationHallucination EvaluationLanguage ModelingLanguage Modelling+2

DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine

2024-11-14 · Jean Seo, Jongwon Lim, Dongjun Jang, Hyopil Shin

We introduce DAHL, a benchmark dataset and automated evaluation system designed to assess hallucination in long-form text generation, specifically within the biomedical domain. Our benchmark dataset, meticulously curated…

FormHallucinationHallucination EvaluationMultiple-choice+1

DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark

2024-11-05 · Haodong Li, Haicheng Qu, Xiaofeng Zhang

With the rapid development of large vision language models (LVLMs), these models have shown excellent results in various multimodal tasks. Since LVLMs are prone to hallucinations and there are currently few datasets and …

Data AugmentationHallucinationHallucination Evaluation

Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models

2024-10-30 · Junjie Wu, Tsz Ting Chung, Kai Chen, Dit-yan Yeung

Despite the outstanding performance in vision-language reasoning, Large Vision-Language Models (LVLMs) might generate hallucinated contents that do not exist in the given image. Most existing LVLM hallucination benchmark…

HallucinationHallucination EvaluationObjectObject Hallucination+2

A Survey of Hallucination in Large Visual Language Models

2024-10-20 · Wei Lan, WenYi Chen, Qingfeng Chen, Shirui Pan 외

The Large Visual Language Models (LVLMs) enhances user interaction and enriches user experience by integrating visual modality on the basis of the Large Language Models (LLMs). It has demonstrated their powerful informat…

HallucinationHallucination EvaluationSurvey

LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models

2024-10-13 · Han Qiu, Jiaxing Huang, Peng Gao, Qin Qi 외

Hallucination, a phenomenon where multimodal large language models~(MLLMs) tend to generate textual responses that are plausible but unaligned with the image, has become one major hurdle in various MLLM-related applicati…

HallucinationHallucination EvaluationMultiple-choice

TLDR: Token-Level Detective Reward Model for Large Vision Language Models

2024-10-07 · Deqing Fu, Tong Xiao, Rui Wang, Wang Zhu 외

Although reward models have been successful in improving multimodal large language models, the reward models themselves remain brutal and contain minimal information. Notably, existing reward models only mimic human anno…

HallucinationHallucination Evaluation

Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization

2024-09-22 · Minyi Zhao, Jie Wang, Zhaoyang Li, Jiyuan Zhang 외

Recent studies have shown that Vision Language Large Models (VLLMs) may output content not relevant to the input images. This problem, called the hallucination phenomenon, undoubtedly degrades VLLM performance. Therefore…

HallucinationHallucination EvaluationImage CaptioningResponse Generation

FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs

2024-09-20 · Bowen Yan, Zhengsong Zhang, Liqiang Jing, Eftekhar Hossain 외

The rapid development of Large Vision-Language Models (LVLMs) often comes with widespread hallucination issues, making cost-effective and comprehensive assessments increasingly vital. Current approaches mainly rely on co…

HallucinationHallucination Evaluation
1–20 / 49 다음 →