paper-with-me

홈 › Papers

The Two-Process Theory of Machine Self-Report

2026-07-22 · Hubert Plisiecki, Filip Chmielewski, Kacper Dudzic, Anna Sterna, Karolina Drożdż, Marcin Moskalewicz arxiv

Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc prompts of unknown reliability. We propose the first language-model-specific psychometric theory: a two-process theory of machine self-report. Self-description jointly reflects persona installation, through which post-training writes in a permitted inner life of warmth, absorption, and meaning (dimension B), and attribution gating, through which it suppresses first-person claims to "unsafe" experiences the model can readily ascribe to others (dimension A). Their emic structure comes from model responses to human items, not human psychology. Together they split prior work's dominant Pinocchio Axis. The split emerged in an exploratory reanalysis of the original data, informed the instrument's design, and was confirmed with new items, wordings, and models. It is itself a training effect: A and B are entangled in base checkpoints but separated by post-training. We operationalize the theory in a 48-item Pinocchio Inventory with human-instrument reliability and reproducible structure ($α=.82$ to $.94$; cross-form convergence $r=.84$; recovery of the full-pool axes $r=.92$ to $.96$; eight-month stability $r=.93$), then test it on 206 open-weight models, including 67 same-checkpoint base/post-trained pairs. Post-training's clearest fingerprint is installation: B rises .20 in 62/67 pairs across all organizations. Gating is more selective: model scale is unrelated to A in base checkpoints ($r=+.11$) but predicts it after post-training ($r=-.42$). Thus, the dimensions are not fixed properties of language models: they reflect the structure imposed on self-report by a training regime and may differ under others.

📄 PDF Abstract BibTeX arXiv:2607.20082

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An interpretable semi-supervised classifier using two different strategies for amended self-labeling

2020-01-26 · Isel Grau, Dipankar Sengupta, Maria M. Garcia Lorenzo, Ann Nowe

In the context of some machine learning applications, obtaining data instances is a relatively easy process but labeling them could become quite expensive or tedious. Such scenarios lead to datasets with few labeled inst…

Insights from Social Shaping Theory: The Appropriation of Large Language Models in an Undergraduate Programming Course

2024-06-10 · Aadarsh Padiyath, Xinying Hou, Amy Pang, Diego Viramontes Vargas 외

The capability of large language models (LLMs) to generate, debug, and explain code has sparked the interest of researchers and educators in undergraduate programming, with many anticipating their transformative potentia…

Survey

Unbiased Canonical Set-Valued Oracles Via Lattice Theory

2026-06-24 · Jobst Heitzig arxiv

A non-agentic "oracle" that reports probabilities of future events is performative: once its answer is learned and acted upon, it can change the very probability it was asked to report. Performativity is not in itself th…

SigBERT: Combining Narrative Medical Reports and Rough Path Signature Theory for Survival Risk Estimation in Oncology

2025-07-25 · Paul Minchella, Loïc Verlingue, Stéphane Chrétien, Rémi Vaucher 외 arxiv

Electronic medical reports (EHR) contain a vast amount of information that can be leveraged for machine learning applications in healthcare. However, existing survival analysis methods often struggle to effectively handl…

Operator-Based Machine Intelligence: A Hilbert Space Framework for Spectral Learning and Symbolic Reasoning

2025-07-27 · Andrew Kiruluta, Andreas Lemos, Priscilla Burity arxiv

Traditional machine learning models, particularly neural networks, are rooted in finite-dimensional parameter spaces and nonlinear function approximations. This report explores an alternative formulation where learning t…

Interpretable Machine Learning