paper-with-me

홈 › Papers

Representational and Behavioral Stability of Truth in Large Language Models

2025-11-24 · Samantha Dies, Courtney Maynard, Germans Savcisens, Tina Eliassi-Rad arxiv

Large language models (LLMs) are increasingly used as information sources, yet small changes in semantic framing can destabilize their truth judgments. We propose P-StaT (Perturbation Stability of Truth), an evaluation framework for testing belief stability under controlled semantic perturbations in representational and behavioral settings via probing and zero-shot prompting. Across sixteen open-source LLMs and three domains, we compare perturbations involving epistemically familiar Neither statements drawn from well-known fictional contexts (Fictional) to those involving unfamiliar Neither statements not seen in training data (Synthetic). We find a consistent stability hierarchy: Synthetic content aligns closely with factual representations and induces the largest retractions of previously held beliefs, producing up to $32.7\%$ retractions in representational evaluations and up to $36.3\%$ in behavioral evaluations. By contrast, Fictional content is more representationally distinct and comparatively stable. Together, these results suggest that epistemic familiarity is a robust signal across instantiations of belief stability under semantic reframing, complementing accuracy-based factuality evaluation with a notion of epistemic robustness.

📄 PDF Abstract BibTeX arXiv:2511.19166

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Geometric Stability of Neural Population Codes: Regional Variation, Behavioral Relevance, and Circuit Dependence

2026-06-28 · Prashant C. Raju arxiv

Current models of representational reliability in neural populations focus on temporal stability: whether population centroids are preserved across sessions and days. This framing leaves a fundamental question unanswered…

Cat, Rat, Meow: On the Alignment of Language Model and Human Term-Similarity Judgments

2025-04-10 · Lorenz Linhardt, Tom Neuhäuser, Lenka Tětková, Oliver Eberle

Small and mid-sized generative language models have gained increasing attention. Their size and availability make them amenable to being analyzed at a behavioral as well as a representational level, allowing investigatio…

Language ModelingLanguage ModellingTriplet

Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions

2026-06-25 · Abla Bedoui, Ashley L. Greene, Mohammed Cherkaoui arxiv

Large language models (LLMs) are increasingly being integrated into mental health support tools and other psychologically sensitive conversational applications. In such settings, behavioral stability and consistency are …

When Role-playing, Do Models Believe What They Say?

2026-06-09 · Benjamin Sturgeon, David Africa, Sid Black arxiv

Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adoption is fundamental to how language models behave, with models selectin…

Representational Curvature Modulates Behavioral Uncertainty in Large Language Models

2026-04-27 · Jack King, Evelina Fedorenko, Eghbal A. Hosseini arxiv

In autoregressive large language models (LLMs), temporal straightening offers an account of how the next-token prediction objective shapes representations. Models learn to progressively straighten the representational tr…