paper-with-me

홈 › Papers

Beyond Mimicry: Preference Coherence in LLMs

2025-11-17 · Luhan Mikaelson, Derek Shiller, Hayley Clatterbuck arxiv

We investigate whether large language models exhibit genuine preference structures by testing their responses to AI-specific trade-offs involving GPU reduction, capability restrictions, shutdown, deletion, oversight, and leisure time allocation. Analyzing eight state-of-the-art models across 48 model-category combinations using logistic regression and behavioral classification, we find that 23 combinations (47.9%) demonstrated statistically significant relationships between scenario intensity and choice patterns, with 15 (31.3%) exhibiting within-range switching points. However, only 5 combinations (10.4%) demonstrate meaningful preference coherence through adaptive or threshold-based behavior, while 26 (54.2%) show no detectable trade-off behavior. The observed patterns can be explained by three distinct decision-making architectures: comprehensive trade-off systems, selective trigger mechanisms, and no stable decision-making paradigm. Testing an instrumental hypothesis through temporal horizon manipulation reveals paradoxical patterns inconsistent with pure strategic optimization. The prevalence of unstable transitions (45.8%) and stimulus-specific sensitivities suggests current AI systems lack unified preference structures, raising concerns about deployment in contexts requiring complex value trade-offs.

📄 PDF Abstract BibTeX arXiv:2511.13630

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Adaptation Paradox: Agency vs. Mimicry in Companion Chatbots

2025-09-16 · T. James Brandt, Cecilia Xi Wang arxiv

Generative AI powers a growing wave of companion chatbots, yet principles for fostering genuine connection remain unsettled. We test two routes: visible user authorship versus covert language-style mimicry. In a preregis…

Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning

2026-01-07 · Soheil Zibakhsh Shabgahi, Pedram Aghazadeh, Farinaz Koushanfar arxiv

Supervised Fine-Tuning (SFT) is a standard approach for injecting domain knowledge into Large Language Models (LLMs). However, relying on validation perplexity to monitor training is often insufficient, as it confounds s…

Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity

2022-09-24 · Gabriel Simmons

Large Language Models (LLMs) have demonstrated impressive capabilities in generating fluent text, as well as tendencies to reproduce undesirable social biases. This study investigates whether LLMs reproduce the moral bia…

Language Modelling

Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization

2026-04-24 · Weixu Zhang, Ye Yuan, Changjiang Han, Yuxing Tian 외 arxiv

Large Language Models (LLMs) exhibit strong implicit personalization ability, yet most existing approaches treat this behavior as a black box, relying on prompt engineering or fine tuning on user data. In this work, we a…

Prompt Engineering

MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT Distillation

2026-01-07 · Jin Cui, Jiaqi Guo, Jiepeng Zhou, Ruixuan Yang 외 arxiv

While Large Language Models (LLMs) have emerged with remarkable capabilities in complex tasks through Chain-of-Thought reasoning, practical resource constraints have sparked interest in transferring these abilities to sm…

Domain Generalization