paper-with-me

홈 › Papers

When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

2022-12-20 · Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi

Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the limitations of relying solely on their parameters to encode a wealth of world knowledge. This paper aims to understand LMs' strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments of 10 models and 4 augmentation methods on PopQA, our new open-domain QA dataset with 14k questions. We find that LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the long tail. We then show that retrieval-augmented LMs largely outperform orders of magnitude larger LMs, while unassisted LMs remain competitive in questions about high-popularity entities. Based on those findings, we devise a simple, yet effective, method for powerful and efficient retrieval-augmented LMs, which retrieves non-parametric memories only when necessary. Experimental results show that this significantly improves models' performance while reducing the inference costs.

📄 PDF Abstract BibTeX arXiv:2212.10511

Code (1)

alextmallen/adaptive-retrieval 공식 구현 pytorch

Tasks

Knowledge ProbingMemorizationRetrievalWorld Knowledge

Similar Papers 제목 키워드 기반

Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning

2024-10-22 · Zongmeng Zhang, Yufeng Shi, Jinhua Zhu, Wengang Zhou 외

Trustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite bein…

RetrievalRetrieval-augmented Generation

Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension

2024-02-28 · Fan Yin, Jayanth Srinivasa, Kai-Wei Chang

We study how to characterize and predict the truthfulness of texts generated from large language models (LLMs), which serves as a crucial step in building trust between humans and LLMs. Although several approaches based …

Language ModelingLanguage ModellingLarge Language ModelQuestion Answering

When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation

2026-08-17 · Haolin Jin, Pengyue Yang, Huaming Chen arxiv

Retrieval-augmented generation (RAG) improves large language models by grounding generation in external evidence, but it also introduces a source trust problem: retrieved context may be useful, irrelevant, or even mislea…

TrustMargin: Training-Free Arbitration between Parametric Memory and Retrieved Evidence in Large Language Models

2026-06-07 · Jingyan Xu, Hong Shi, Yi Shan, Penghui Liu 외 arxiv

Large language models answer knowledge-intensive questions using both parametric memory and retrieved evidence, but neither source is uniformly reliable. Retrieval can fill knowledge gaps, yet distracting passages may ov…

Simulations evaluating resampling methods for causal discovery: ensemble performance and calibration

2019-10-04 · Erich Kummerfeld, Alexander Rix

Causal discovery can be a powerful tool for investigating causality when a system can be observed but is inaccessible to experiments in practice. Despite this, it is rarely used in any scientific or medical fields. One o…

Causal Discovery