paper-with-me

홈 › Papers

Improving Factuality with Explicit Working Memory

2024-12-24 · Mingda Chen, Yang Li, Karthik Padthe, Rulin Shao, Alicia Sun, Luke Zettlemoyer, Gargi Gosh, Wen-tau Yih

Large language models can generate factually inaccurate content, a problem known as hallucination. Recent works have built upon retrieved-augmented generation to improve factuality through iterative prompting but these methods are limited by the traditional RAG design. To address these challenges, we introduce EWE (Explicit Working Memory), a novel approach that enhances factuality in long-form text generation by integrating a working memory that receives real-time feedback from external resources. The memory is refreshed based on online fact-checking and retrieval feedback, allowing EWE to rectify false claims during the generation process and ensure more accurate and reliable outputs. Our experiments demonstrate that Ewe outperforms strong baselines on four fact-seeking long-form generation datasets, increasing the factuality metric, VeriScore, by 2 to 10 points absolute without sacrificing the helpfulness of the responses. Further analysis reveals that the design of rules for memory updates, configurations of memory units, and the quality of the retrieval datastore are crucial factors for influencing model performance.

📄 PDF Abstract BibTeX arXiv:2412.18069

Code (0)

등록된 구현이 없습니다.

Tasks

Fact CheckingHallucinationRAGRetrievalText Generation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

A Simple Reservoir Model of Working Memory with Real Values

2018-06-18 · Anthony Strock, Nicolas Rougier, Xavier Hinaut

The prefrontal cortex is known to be involved in many high-level cognitive functions, in particular, working memory. Here, we study to what extent a group of randomly connected units (namely an Echo State Network, ESN) c…

Complexity of Symbolic Representation in Working Memory of Transformer Correlates with the Complexity of a Task

2024-06-20 · Alsu Sagirova, Mikhail Burtsev

Even though Transformers are extensively used for Natural Language Processing tasks, especially for machine translation, they lack an explicit memory to store key concepts of processed texts. This paper explores the prop…

DecoderDiversityMachine TranslationTranslation

Preliminary Evidence -- Diagnosed Alzheimer's Disease But Not MCI Affects Working Memory Capacity - 0.7 of 2.7 Memory Slots is Lost

2016-03-24

Recently it was shown explicitly that free recall consists of two stages: the first few recalls empty working memory (narrowly defined) and a second stage, a reactivation stage, concludes the recall (Tarnow, 2015). It wa…

Management

From implicit learning to explicit representations

2022-04-05 · Naomi Chaix-Eichel, Snigdha Dagar, Quentin Lanneau, Karen Sobriel 외

Using the reservoir computing framework, we demonstrate how a simple model can solve an alternation task without an explicit working memory. To do so, a simple bot equipped with sensors navigates inside a 8-shaped maze a…

Working memory capacity and gender

2017-03-21

The working memory capacity (WMC) of 400 Russian college students was measured using the Tarnow Unchunkable Test [2] which tests WMC alone without requiring explicit working memory operations. We found small-sized WMC di…

Math