paper-with-me

Papers

Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning

2026-03-01 · Hamed Damirchi, Ignacio Meza De la Jara, Ehsan Abbasnejad, Afshar Shamsi, Zhen Zhang, Javen Shi arxiv

Existing explainability methods for Large Language Models (LLMs) typically treat hidden states as static points in activation space, assuming that correct and incorrect inferences can be separated using representations from an individual layer. However, these activations are saturated with polysemantic features, leading to linear probes learning surface-level lexical patterns rather than underlying reasoning structures. We introduce Truth as a Trajectory (TaT), which models the transformer inference as an unfolded trajectory of iterative refinements, shifting analysis from static activations to layer-wise geometric displacement. By analyzing displacement of representations across layers, TaT uncovers geometric invariants that distinguish valid reasoning from spurious behavior. We evaluate TaT across dense and Mixture-of-Experts (MoE) architectures on benchmarks spanning commonsense reasoning, question answering, and toxicity detection. Without access to the activations themselves and using only changes in activations across layers, we show that TaT effectively mitigates reliance on static lexical confounds, outperforming conventional probing, and establishes trajectory analysis as a complementary perspective on LLM explainability.

📄 PDF Abstract BibTeX arXiv:2603.01326

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

2024-10-03 · Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart 외

Large language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have demonstrated that LLMs' internal states…

TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space

2024-02-27 · Shaolei Zhang, Tian Yu, Yang Feng

Large Language Models (LLMs) sometimes suffer from producing hallucinations, especially LLMs may generate untruthful responses despite knowing the correct knowledge. Activating the truthfulness within LLM is the key to f…

Contrastive LearningHallucinationHallucination EvaluationLanguage Modelling+4

Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness

2025-10-10 · Chi Seng Cheang, Hou Pong Chan, Wenxuan Zhang, Yang Deng arxiv

Recent work suggests that LLMs "know what they don't know", positing that hallucinated and factually correct outputs arise from distinct internal processes and can therefore be distinguished using internal signals. Howev…

Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations

2026-01-12 · Wen Luo, Guangyue Peng, Wei Li, Shaohang Wei 외 arxiv

Despite their impressive capabilities, large language models (LLMs) frequently generate hallucinations. Previous work shows that their internal states encode rich signals of truthfulness, yet the origins and mechanisms o…

Overlapping neural representations for the position of visible and imagined objects

2020-11-11

Humans can covertly track the position of an object, even if the object is temporarily occluded. What are the neural mechanisms underlying our capacity to track moving objects when there is no physical stimulus for the b…

EEGElectroencephalogram (EEG)ObjectObject Tracking+1