Strategic Behavior under Context Misalignment
We study the behavioral implications of Rationality and Common Strong Belief in Rationality (RCSBR) with contextual assumptions allowing players to entertain misaligned beliefs, i.e., players can hold beliefs concerning their opponents' beliefs where there is no opponent holding those very beliefs. Taking the analysts' perspective, we distinguish the infinite hierarchies of beliefs actually held by players ("real types") from those that are a byproduct of players' hierarchies ("imaginary types") by introducing the notion of separating type structure. We characterize the behavioral implications of RCSBR for the real types across all separating type structures via a family of subsets of Full Strong Best-Reply Sets of Battigalli & Friedenberg (2012). By allowing misalignment, in dynamic games we can obtain behavioral predictions inconsistent with RCSBR (in the standard framework), contrary to the case of belief-based analyses for static games--a difference due to the dichotomy "non-monotonic vs. monotonic" reasoning.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior
The alignment problem refers to concerns regarding powerful intelligences, ensuring compatibility with human preferences and values as capabilities increase. Current large language models (LLMs) show misaligned behaviors…
Probing the Misaligned Thinking Process of Language Models
Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-stakes settings, it is critical to reliably…
Sycophancy Towards Researchers Drives Performative Misalignment
The increasing situational awareness of language models raises safety concerns: models might be aware when they are evaluated, and adjust their behavior to evade monitoring and resist modification, e.g., pretending to be…
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reinforcement Learning from Human Feedback (RLHF) and related alignment paradigms have become central to steering large language models (LLMs) and multimodal large language models (MLLMs) toward human-preferred behaviors…
Reinforcement LearningThe Law-Following AI Framework: Legal Foundations and Technical Constraints. Legal Analogues for AI Actorship and technical feasibility of Law Alignment
This paper critically evaluates the "Law-Following AI" (LFAI) framework proposed by O'Keefe et al. (2025), which seeks to embed legal compliance as a superordinate design objective for advanced AI agents and enable them …