paper-with-me

홈 › Papers

Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability

2026-06-08 · Emirhan Bilgiç, Baptiste Caramiaux, Zhi Yan, Gianni Franchi arxiv

As Vision-Language Models are increasingly deployed in safety-critical applications, the trustworthiness of their explanations becomes crucial. Explainable AI (XAI) methods for Vision-Language Models often suffer from semantic hallucination, where attribution maps highlight prominent image regions even when prompted with incorrect text descriptions (e.g., highlighting a dog when prompted ``cat''). Although this problem is widespread, a formal mathematical analysis of XAI methods and CLIP embeddings is largely missing in the literature. We demonstrate that this phenomenon is not specific to a single architecture but is a fundamental consequence of Linear Semantic Leakage in high-dimensional embedding spaces. We propose a unified theoretical framework, Linear Semantic Attribution (LSA), which generalizes across discriminative methods. We introduce OSP, a geometric intervention that utilizes the residual property of OMP to disentangle unique semantic signals from shared concepts. We prove theoretically and demonstrate empirically that OSP minimizes hallucination by orthogonalizing the query vector against distractor concepts, rendering the attribution model blind to shared features while preserving fidelity for correct prompts. Our code is available at: https://github.com/emirhanbilgic/Orthogonal-Semantic-Projection

📄 PDF Abstract BibTeX arXiv:2606.14758

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word Embeddings

2020-06-30 · EMNLP 2021 11 · Sunipa Dev, Tao Li, Jeff M. Phillips, Vivek Srikumar

Language representations are known to carry stereotypical biases and, as a result, lead to biased predictions in downstream tasks. While existing methods are effective at mitigating biases by linear projection, such meth…

Word Embeddings

HARP: Hallucination Detection via Reasoning Subspace Projection

2025-09-15 · Junjie Hu, Gang Tu, ShengYu Cheng, Jinxin Li 외 arxiv

Hallucinations in Large Language Models (LLMs) pose a major barrier to their reliable use in critical decision-making. Although existing hallucination detection methods have improved accuracy, they still struggle with di…

ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability

2024-10-15 · Zhongxiang Sun, Xiaoxue Zang, Kai Zheng, Jun Xu 외

Retrieval-Augmented Generation (RAG) models are designed to incorporate external knowledge, reducing hallucinations caused by insufficient parametric (internal) knowledge. However, even with accurate and relevant retriev…

HallucinationRAGRetrievalRetrieval-augmented Generation

Hallucinations as Orthogonal Noise: Inference-Time Manifold Alignment via Dynamic Contextual Orthogonalization

2026-06-02 · Mingkuan Zhao, Wentao Hu, Tianchen Huang, Yuheng Min 외 arxiv

Hallucination in Large Language Models (LLMs), characterized by the generation of content inconsistent with contextual facts or logical constraints -- remains a persistent challenge for reliable deployment. In this work,…

Disentangling Polysemantic Channels in Convolutional Neural Networks

2025-04-17 · Robin Hesse, Jonas Fischer, Simone Schaub-Meyer, Stefan Roth

Mechanistic interpretability is concerned with analyzing individual components in a (convolutional) neural network (CNN) and how they form larger circuits representing decision mechanisms. These investigations are challe…