paper-with-me

Papers

Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks

2023-06-30 · B. A. Levinstein, Daniel A. Herrmann

We consider the questions of whether or not large language models (LLMs) have beliefs, and, if they do, how we might measure them. First, we evaluate two existing approaches, one due to Azaria and Mitchell (2023) and the other to Burns et al. (2022). We provide empirical results that show that these methods fail to generalize in very basic ways. We then argue that, even if LLMs have beliefs, these methods are unlikely to be successful for conceptual reasons. Thus, there is still no lie-detector for LLMs. After describing our empirical results we take a step back and consider whether or not we should expect LLMs to have something like beliefs in the first place. We consider some recent arguments aiming to show that LLMs cannot have beliefs. We show that these arguments are misguided. We provide a more productive framing of questions surrounding the status of beliefs in LLMs, and highlight the empirical nature of the problem. We conclude by suggesting some concrete paths for future work.

📄 PDF Abstract BibTeX arXiv:2307.00175

Code (1)

balevinstein/probes 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs

2025-02-27 · Linyang He, Ercong Nie, Sukru Samet Dindar, Arsalan Firoozi 외

We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct…

Knowledge Distillation

COPEN: Probing Conceptual Knowledge in Pre-trained Language Models

2022-11-08 · Hao Peng, Xiaozhi Wang, Shengding Hu, Hailong Jin 외

Conceptual knowledge is fundamental to human cognition and knowledge bases. However, existing knowledge probing works only focus on evaluating factual knowledge of pre-trained language models (PLMs) and ignore conceptual…

Knowledge Probing

Probing Neural Language Models for Human Tacit Assumptions

2020-04-10 · Nathaniel Weir, Adam Poliak, Benjamin Van Durme

Humans carry stereotypic tacit assumptions (STAs) (Prince, 1978), or propositional beliefs about generic concepts. Such associations are crucial for understanding natural language. We construct a diagnostic set of word p…

Diagnostic

Probing the Limits of the Lie Detector Approach to LLM Deception

2026-02-16 · Tom-Felix Berger arxiv

Mechanistic approaches to deception in large language models (LLMs) often rely on "lie detectors", that is, truth probes trained to identify internal representations of model outputs as false. The lie detector approach t…

Ranking Entities along Conceptual Space Dimensions with LLMs: An Analysis of Fine-Tuning Strategies

2024-02-23 · Nitesh Kumar, Usashi Chatterjee, Steven Schockaert

Conceptual spaces represent entities in terms of their primitive semantic features. Such representations are highly valuable but they are notoriously difficult to learn, especially when it comes to modelling perceptual a…