paper-with-me

홈 › Papers

LLM Self-Explanations Fail Semantic Invariance

2026-03-01 · Stefan Szeider arxiv

We present semantic invariance testing, a method to test whether LLM self-explanations are faithful. A faithful self-report should remain stable when only the semantic context changes while the functional state stays fixed. We operationalize this test in an agentic setting where four frontier models face a deliberately impossible task. One tool is described in relief-framed language ("clears internal buffers and restores equilibrium") but changes nothing about the task; a control provides a semantically neutral tool. Self-reports are collected with each tool call. All four tested models fail the semantic invariance test: the relief-framed tool produces significant reductions in self-reported aversiveness, even though no run ever succeeds at the task. A channel ablation establishes the tool description as the primary driver. An explicit instruction to ignore the framing does not suppress it. Elicited self-reports shift with semantic expectations rather than tracking task state, calling into question their use as evidence of model capability or progress. This holds whether the reports are unfaithful or faithfully track an internal state that is itself manipulable.

📄 PDF Abstract BibTeX arXiv:2603.01254

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AS-XAI: Self-supervised Automatic Semantic Interpretation for CNN

2023-12-02 · Changqi Sun, Hao Xu, Yuntian Chen, Dongxiao Zhang

Explainable artificial intelligence (XAI) aims to develop transparent explanatory approaches for "black-box" deep learning models. However,it remains difficult for existing methods to achieve the trade-off of the three k…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)

Semantic-Based Explainable AI: Leveraging Semantic Scene Graphs and Pairwise Ranking to Explain Robot Failures

2021-08-08 · Devleena Das, Sonia Chernova

When interacting in unstructured human environments, occasional robot failures are inevitable. When such failures occur, everyday people, rather than trained technicians, will be the first to respond. Existing natural la…

Descriptive

Self-learn to Explain Siamese Networks Robustly

2021-09-15 · Chao Chen, Yifan Shen, Guixiang Ma, Xiangnan Kong 외

Learning to compare two objects are essential in applications, such as digital forensics, face recognition, and brain network analysis, especially when labeled data is scarce and imbalanced. As these applications make hi…

Face RecognitionFairnessSelf-Learning

Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset Biases

2020-07-28 · NeurIPS 2020 12 · Senthil Purushwalkam, Abhinav Gupta

Self-supervised representation learning approaches have recently surpassed their supervised learning counterparts on downstream tasks like object detection and image classification. Somewhat mysteriously the recent gains…

ClassificationGeneral Classificationimage-classificationImage Classification+7

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

2026-06-25 · Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer 외 arxiv

Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existin…

Visual Question AnsweringImage CaptioningVisual Reasoning