paper-with-me

홈 › Papers

KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions

2024-07-08 · Yanxu Zhu, Jinlin Xiao, Yuhang Wang, Jitao Sang

Recent studies have demonstrated that large language models (LLMs) are susceptible to being misled by false premise questions (FPQs), leading to errors in factual knowledge, know as factuality hallucination. Existing benchmarks that assess this vulnerability primarily rely on manual construction, resulting in limited scale and lack of scalability. In this work, we introduce an automated, scalable pipeline to create FPQs based on knowledge graphs (KGs). The first step is modifying true triplets extracted from KGs to create false premises. Subsequently, utilizing the state-of-the-art capabilities of GPTs, we generate semantically rich FPQs. Based on the proposed method, we present a comprehensive benchmark, the Knowledge Graph-based False Premise Questions (KG-FPQ), which contains approximately 178k FPQs across three knowledge domains, at six levels of confusability, and in two task formats. Using KG-FPQ, we conduct extensive evaluations on several representative LLMs and provide valuable insights. The KG-FPQ dataset and code are available at~https://github.com/yanxuzhu/KG-FPQ.

📄 PDF Abstract BibTeX arXiv:2407.05868

Code (1)

yanxuzhu/kg-fpq 공식 구현

Tasks

HallucinationKnowledge Graphs

Similar Papers 제목 키워드 기반

WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries

2024-07-24 · Wenting Zhao, Tanya Goyal, Yu Ying Chiu, Liwei Jiang 외

While hallucinations of large language models (LLMs) prevail as a major challenge, existing evaluation benchmarks on factuality do not cover the diverse domains of knowledge that the real-world users of LLMs seek informa…

ChatbotFormHallucinationRetrieval

The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework

2025-10-29 · Aakriti Shah, Thai Le arxiv

Unlearning in large language models (LLMs) is crucial for managing sensitive data and correcting misinformation, yet evaluating its effectiveness remains an open problem. We investigate whether persuasive prompting can r…

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning

2025-07-25 · Shengyuan Wang, Jie Feng, Tianhui Liu, Dan Pei 외 arxiv

Large language models (LLMs) possess extensive world knowledge, including geospatial knowledge, which has been successfully applied to various geospatial tasks such as mobility prediction and social indicator prediction.…

General KnowledgeKnowledge Graphs

AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains

2026-01-21 · Adam Szelestey, Sofie van Engelen, Tianhao Huang, Justin Snelders 외 arxiv

Hallucination in large language models (LLMs) remains an acute concern, contributing to the spread of misinformation and diminished public trust, particularly in high-risk domains. Among hallucination types, factuality i…

Adversarial Attack

Evaluating LLMs' Assessment of Mixed-Context Hallucination Through the Lens of Summarization

2025-03-03 · Siya Qi, Rui Cao, Yulan He, Zheng Yuan

With the rapid development of large language models (LLMs), LLM-as-a-judge has emerged as a widely adopted approach for text quality evaluation, including hallucination evaluation. While previous studies have focused exc…

HallucinationHallucination Evaluation