paper-with-me

홈 › Papers

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis

2025-05-20 · Haoming Huang, Yibo Yan, Jiahao Huo, Xin Zou, Xinfeng Li, Kun Wang, Xuming Hu

Large Language Models (LLMs), despite their remarkable capabilities, are hampered by hallucinations. A particularly challenging variant, knowledge overshadowing, occurs when one piece of activated knowledge inadvertently masks another relevant piece, leading to erroneous outputs even with high-quality training data. Current understanding of overshadowing is largely confined to inference-time observations, lacking deep insights into its origins and internal mechanisms during model training. Therefore, we introduce PhantomCircuit, a novel framework designed to comprehensively analyze and detect knowledge overshadowing. By innovatively employing knowledge circuit analysis, PhantomCircuit dissects the internal workings of attention heads, tracing how competing knowledge pathways contribute to the overshadowing phenomenon and its evolution throughout the training process. Extensive experiments demonstrate PhantomCircuit's effectiveness in identifying such instances, offering novel insights into this elusive hallucination and providing the research community with a new methodological lens for its potential mitigation.

📄 PDF Abstract BibTeX arXiv:2505.14406

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucination

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination

2025-02-22 · Yuji Zhang, Sha Li, Cheng Qian, Jiateng Liu 외

Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which error generation continues despite high-qua…

HallucinationText Generation

Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models

2024-07-10 · Yuji Zhang, Sha Li, Jiateng Liu, Pengfei Yu 외

Hallucination is often regarded as a major impediment for using large language models (LLMs), especially for knowledge-intensive tasks. Even when the training corpus consists solely of true statements, language models st…

HallucinationLanguage ModelingLanguage Modelling

$\varepsilon$ KÚ <MASK>: Integrating Yorùbá cultural greetings into machine translation

2023-03-31 · Idris Akinade, Jesujoba Alabi, David Adelani, Clement Odoje 외

This paper investigates the performance of massively multilingual neural machine translation (NMT) systems in translating Yor\`ub\'a greetings ($\varepsilon$ k\'u [MASK]), which are a big part of Yor\`ub\'a language and …

Cultural Vocal Bursts Intensity PredictionMachine TranslationNMTTranslation

ActiShade: Activating Overshadowed Knowledge to Guide Multi-Hop Reasoning in Large Language Models

2026-01-12 · Huipeng Ma, Luan Zhang, Dandan Song, Linmei Hu 외 arxiv

In multi-hop reasoning, multi-round retrieval-augmented generation (RAG) methods typically rely on LLM-generated content as the retrieval query. However, these approaches are inherently vulnerable to knowledge overshadow…

Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization

2026-03-13 · Zequn Liu, Kehan Wu, Shufang Xie, Zekun Guo 외 arxiv

Emerging reasoning models hold promise for automating scientific discovery. However, their training is hindered by a critical supervision gap: experimental outcomes are abundant, whereas intermediate reasoning steps are …

Drug Discovery