paper-with-me

홈 › Papers

When Truthful Representations Flip Under Deceptive Instructions?

2025-07-29 · Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng, Haotian Yu, Xiaotian Han, Pan Li arxiv

Large language models (LLMs) tend to follow maliciously crafted instructions to generate deceptive responses, posing safety challenges. How deceptive instructions alter the internal representations of LLM compared to truthful ones remains poorly understood beyond output analysis. To bridge this gap, we investigate when and how these representations ``flip'', such as from truthful to deceptive, under deceptive versus truthful/neutral instructions. Analyzing the internal representations of Llama-3.1-8B-Instruct and Gemma-2-9B-Instruct on a factual verification task, we find the model's instructed True/False output is predictable via linear probes across all conditions based on the internal representation. Further, we use Sparse Autoencoders (SAEs) to show that the Deceptive instructions induce significant representational shifts compared to Truthful/Neutral representations (which are similar), concentrated in early-to-mid layers and detectable even on complex datasets. We also identify specific SAE features highly sensitive to deceptive instruction and use targeted visualizations to confirm distinct truthful/deceptive representational subspaces. % Our analysis pinpoints layer-wise and feature-level correlates of instructed dishonesty, offering insights for LLM detection and control. Our findings expose feature- and layer-level signatures of deception, offering new insights for detecting and mitigating instructed dishonesty in LLMs.

📄 PDF Abstract BibTeX arXiv:2507.22149

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs

2025-08-26 · Yao Fu, Xianxuan Long, Runchao Li, Haotian Yu 외 arxiv

Quantization enables efficient deployment of large language models (LLMs) in resource-constrained environments by significantly reducing memory and computation costs. While quantized LLMs often maintain performance on pe…

Logical Reasoning

When lies are mostly truthful: automated verbal deception detection for embedded lies

2025-01-13 · Riccardo Loconte, Bennett Kleinberg

Background: Verbal deception detection research relies on narratives and commonly assumes statements as truthful or deceptive. A more realistic perspective acknowledges that the veracity of statements exists on a continu…

Deception Detection

A Multimodal Dataset for Deception Detection

2014-05-01 · LREC 2014 5 · Ver{\'o}nica P{\'e}rez-Rosas, Rada Mihalcea, Alexis Narvaez, Mihai Burzo

This paper presents the construction of a multimodal dataset for deception detection, including physiological, thermal, and visual responses of human subjects under three deceptive scenarios. We present the experimental …

Deception Detection

Linguistic Cues to Deception and Perceived Deception in Interview Dialogues

2018-06-01 · NAACL 2018 6 · Sarah Ita Levitan, Angel Maredia, Julia Hirschberg

We explore deception detection in interview dialogues. We analyze a set of linguistic features in both truthful and deceptive responses to interview questions. We also study the perception of deception, identifying chara…

BIG-bench Machine LearningDeception DetectionGeneral Classification

Is this hotel review truthful or deceptive? A platform for disinformation detection through computational stylometry

2020-05-01 · LREC 2020 5 · Antonio Pascucci, Raffaele Manna, Ciro Caterino, Vincenzo Masucci 외

In this paper, we present a web service platform for disinformation detection in hotel reviews written in English. The platform relies on a hybrid approach of computational stylometry techniques, machine learning and lin…