paper-with-me

홈 › Papers

Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning

2026-06-14 · Zhenyu Yu arxiv

Instruction-tuned language models can answer the same causal-reasoning question differently after its English variable names are replaced by type-preserving placeholders, although the structural causal model and the gold answer are unchanged. We ask whether this lexical gap reflects information loss in the placeholder view or a misaligned read-out from a representation that still carries answer-relevant content. Vernier uses a paired-view weight update as an instrument and then inspects the mechanism left after the gap closes. In the working regimes, the evidence favours representational misalignment. A variable-name probe becomes more accurate on the placeholder view, and activation patching on Qwen-7B, Qwen-14B, and Llama-3.1-8B shows that the decision-token representation can transfer answer identity between views. The update that realigns the views is counterfactual augmentation over original and placeholder prompts, while the answer-subspace KL mainly sharpens intermediate answer-belief agreement. Success is bounded by model family, scale, and task. CRASS transfer is reliable across Qwen scales and Llama, e-CARE remains weak, and preliminary non-causal rename tasks show a similar qualitative pattern.

📄 PDF Abstract BibTeX arXiv:2606.15733

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Attentional White Bear Effect in Transformer Language Models

2026-05-27 · Rebecca Ramnauth, Brian Scassellati arxiv

Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression reduces internal representation or merely suppresses expression. We i…

Reconstruction Probing

2022-12-21 · Najoung Kim, Jatin Khilnani, Alex Warstadt, Abed Qaddoumi

We propose reconstruction probing, a new analysis method for contextualized representations based on reconstruction probabilities in masked language models (MLMs). This method relies on comparing the reconstruction proba…

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

2026-05-26 · Yifan Jiang, Ruoxi Ning, Sheng Yao, Freda Shi arxiv

Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language models (VLMs) can distinguish useful visual evidence from incidental im…

Diagnosing Catastrophe: Large parts of accuracy loss in continual learning can be accounted for by readout misalignment

2023-10-09 · Daniel Anthes, Sushrut Thorat, Peter König, Tim C. Kietzmann

Unlike primates, training artificial neural networks on changing data distributions leads to a rapid decrease in performance on old tasks. This phenomenon is commonly referred to as catastrophic forgetting. In this paper…

Continual Learning

Probing BERT for German Compound Semantics

2025-05-20 · Filip Miletić, Aaron Schmid, Sabine Schulte im Walde

This paper investigates the extent to which pretrained German BERT encodes knowledge of noun compound semantics. We comprehensively vary combinations of target tokens, layers, and cased vs. uncased models, and evaluate t…