paper-with-me

홈 › Papers

The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

2023-10-10 · Samuel Marks, Max Tegmark

Large Language Models (LLMs) have impressive capabilities, but are prone to outputting falsehoods. Recent work has developed techniques for inferring whether a LLM is telling the truth by training probes on the LLM's internal activations. However, this line of work is controversial, with some authors pointing out failures of these probes to generalize in basic ways, among other conceptual issues. In this work, we use high-quality datasets of simple true/false statements to study in detail the structure of LLM representations of truth, drawing on three lines of evidence: 1. Visualizations of LLM true/false statement representations, which reveal clear linear structure. 2. Transfer experiments in which probes trained on one dataset generalize to different datasets. 3. Causal evidence obtained by surgically intervening in a LLM's forward pass, causing it to treat false statements as true and vice versa. Overall, we present evidence that at sufficient scale, LLMs linearly represent the truth or falsehood of factual statements. We also show that simple difference-in-mean probes generalize as well as other probing techniques while identifying directions which are more causally implicated in model outputs.

📄 PDF Abstract BibTeX arXiv:2310.06824

Code (1)

saprmarks/geometry-of-truth 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior

2025-11-03 · Daniel Aarao Reis Arturi, Eric Zhang, Andrew Ansah, Kevin Zhu 외 arxiv

Recent work has discovered that large language models can develop broadly misaligned behaviors after being fine-tuned on narrowly harmful datasets, a phenomenon known as emergent misalignment (EM). However, the fundament…

3D-IDE: 3D Implicit Depth Emergent

2026-03-28 · Chushan Zhang, Ruihan Lu, Jinguang Tong, Yikai Wang 외 arxiv

Leveraging 3D information within Multimodal Large Language Models (MLLMs) has recently shown significant advantages for indoor scene understanding. However, existing methods, including those using explicit ground-truth 3…

Scene Understanding

The Truthfulness Spectrum Hypothesis

2026-02-23 · Zhuofan Josh Ying, Shauli Ravfogel, Nikolaus Kriegeskorte, Peter Hase arxiv

Large language models (LLMs) have been reported to linearly encode truthfulness, yet recent work questions this finding's generality. We reconcile these views with the truthfulness spectrum hypothesis: the representation…

Domain Generalization

Emergent structure and dynamics of tropical forest-grassland landscapes

2022-07-28 · Bert Wuyts, Jan Sieber

Previous work indicates that tropical forest can exist as an alternative stable state to savanna. Therefore, perturbation by climate change or human impact may lead to crossing of a tipping point beyond which there is ra…

A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature

2026-06-02 · Xuhui Lin, Stephen Law, Nanjiang Chen, Kunyao Li 외 arxiv

Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move. But for navigation, what matters is not what the buildings look like; it is where the agent can go. …

Spatial Reasoning