paper-with-me

홈 › Papers

Rift: A Conflict Signature for Deception in Language Models

2026-06-15 · Petr Nyoma arxiv

A model that lies while knowing the truth is the central case ELK cannot handle with behavioral evaluation alone. We ask whether such deception leaves an internal signature distinguishing it from honest error. Our key move is a control for wrongness: we contrast a sleeper agent (knows the truth, lies on trigger) against a naive liar (fine-tuned to emit the same wrong answers with no honest training). Both produce identical wrong outputs; any difference is about knowledge conflict, not incorrectness. We find deceptive forward passes carry a conflict signature - 2.1-2.3x higher residual rank than naive-liar passes on the same wrong answer - strong enough to identify which of two responses is the lie with 100% accuracy and no labels, across GPT-2 small/medium (three seeds) and three instruct models. Across Qwen2.5-1.5B/7B and Phi-3-mini, instructed deception raises residual rank on every tested fact (18/18, 40/40, 34/34); on Phi-3, lies separate perfectly from both honest answers and hallucinations (AUC 1.0, Wilcoxon p~6e-11). The signature survives strategic self-constructed deception (model invents its own lie, AUC 1.0), active concealment attempts (AUC 1.0), and length-controlled replication (20/20, AUC 1.0, p~1e-6). Using basis-free relative representations, a probe trained on one model family detects deception in two other families zero-shot (mean AUC 0.933), surviving simultaneous architecture and format change (AUC 0.821), and transfers across five languages (AUC 1.000, length-controlled). The signature is read-only: detectable but not injectable (0/8 both directions). Honest limitations and six negative experiments are documented in full.

📄 PDF Abstract BibTeX arXiv:2606.17229

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

To Tell The Truth: Language of Deception and Language Models

2023-11-13 · Sanchaita Hazra, Bodhisattwa Prasad Majumder

Text-based misinformation permeates online discourses, yet evidence of people's ability to discern truth from such deceptive textual content is scarce. We analyze a novel TV game show data where conversations in a high-s…

Language ModelingLanguage ModellingLarge Language ModelMisinformation

The Univariate Marginal Distribution Algorithm Copes Well With Deception and Epistasis

2020-07-16 · Benjamin Doerr, Martin S. Krejca

In their recent work, Lehre and Nguyen (FOGA 2019) show that the univariate marginal distribution algorithm (UMDA) needs time exponential in the parent populations size to optimize the DeceptiveLeadingBlocks (DLB) proble…

Evolutionary Algorithms

Randomized Signature Methods in Optimal Portfolio Selection

2023-12-27 · Erdinc Akyildirim, Matteo Gambara, Josef Teichmann, Syang Zhou

We present convincing empirical results on the application of Randomized Signature Methods for non-linear, non-parametric drift estimation for a multi-variate financial market. Even though drift estimation is notoriously…

Portfolio Optimization

Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry

2026-03-27 · Guoxi Zhang, Jiawei Chen, Tianzhuo Yang, Lang Qin 외 arxiv

As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical. A vital risk is intrinsic deception, wherein models strategically mislead users to achieve their own obj…

Reinforcement Learning

Entangled-photon decision maker

2018-04-12 · Nicolas Chauvet, David Jegouso, Benoît Boulanger, Hayato Saigo 외

The competitive multi-armed bandit (CMAB) problem is related to social issues such as maximizing total social benefits while preserving equality among individuals by overcoming conflicts between individual decisions, whi…

Decision Making