paper-with-me

홈 › Papers

Systematic Assessment of Factual Knowledge in Large Language Models

2023-10-18 · Linhao Luo, Thuy-Trang Vu, Dinh Phung, Gholamreza Haffari

Previous studies have relied on existing question-answering benchmarks to evaluate the knowledge stored in large language models (LLMs). However, this approach has limitations regarding factual knowledge coverage, as it mostly focuses on generic domains which may overlap with the pretraining data. This paper proposes a framework to systematically assess the factual knowledge of LLMs by leveraging knowledge graphs (KGs). Our framework automatically generates a set of questions and expected answers from the facts stored in a given KG, and then evaluates the accuracy of LLMs in answering these questions. We systematically evaluate the state-of-the-art LLMs with KGs in generic and specific domains. The experiment shows that ChatGPT is consistently the top performer across all domains. We also find that LLMs performance depends on the instruction finetuning, domain and question complexity and is prone to adversarial context.

📄 PDF Abstract BibTeX arXiv:2310.11638

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge GraphsQuestion Answering

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
FAVOR+ 설명 없음
Performer Performer is a Transformer architecture which can estimate regular…

Similar Papers 제목 키워드 기반

Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation

2023-05-28 · Magdalena Wysocka, Oskar Wysocki, Maxime Delmas, Vincent Mutel 외

The paper introduces a framework for the evaluation of the encoding of factual scientific knowledge, designed to streamline the manual evaluation process typically conducted by domain experts. Inferring over and extracti…

Specificity

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning

2025-07-25 · Shengyuan Wang, Jie Feng, Tianhui Liu, Dan Pei 외 arxiv

Large language models (LLMs) possess extensive world knowledge, including geospatial knowledge, which has been successfully applied to various geospatial tasks such as mobility prediction and social indicator prediction.…

General KnowledgeKnowledge Graphs

Statistical Knowledge Assessment for Large Language Models

2023-09-21 · NeurIPS 2023 11

Given varying prompts regarding a factoid question, can a large language model (LLM) reliably generate factually correct answers? Existing LLMs may generate distinct responses for different prompts. In this paper, we stu…

Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators

2024-06-19 · Matéo Mahaut, Laura Aina, Paula Czarnowska, Momchil Hardalov 외

Large Language Models (LLMs) tend to be unreliable in the factuality of their answers. To address this problem, NLP researchers have proposed a range of techniques to estimate LLM's confidence over facts. However, due to…

Fact VerificationQuestion Answering

Personality Editing for Language Models through Relevant Knowledge Editing

2025-02-17 · Seojin Hwang, Yumin Kim, Byeongjeong Kim, Hwanhee Lee

Large Language Models (LLMs) play a vital role in applications like conversational agents and content creation, where controlling a model's personality is crucial for maintaining tone, consistency, and engagement. Howeve…

knowledge editing