paper-with-me

Papers

The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?

2026-01-30 · Alexander Hägele, Aryo Pradipta Gema, Henry Sleight, Ethan Perez, Jascha Sohl-Dickstein arxiv

As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task scope. It is therefore important to understand how extremely capable AI models will fail: Will they fail by systematically pursuing goals we do not intend? Or will they fail by being a hot mess, and taking nonsensical actions that do not further any goal? We operationalize this question using a bias-variance decomposition of the errors made by AI models: An AI's \emph{error-incoherence} on a task is measured over test-time randomness as the fraction of its error that stems from variance rather than bias in task outcome. Across all tasks and frontier models we measure, the longer models spend reasoning and taking actions, \emph{the more incoherent} their failures become. Error-incoherence changes with model scale in a way that is experiment dependent. However, in several settings, larger, more capable models are more incoherent than smaller models. Consequently, scale alone seems unlikely to eliminate error-incoherence. Instead, as more capable AIs pursue harder tasks, requiring more sequential action and thought, our results predict failures to be accompanied by more incoherent behavior. This suggests a future where AIs sometimes cause industrial accidents (due to unpredictable misbehavior), but are less likely to exhibit consistent pursuit of a misaligned goal. This increases the relative importance of alignment research targeting reward hacking or goal misspecification.

📄 PDF Abstract BibTeX arXiv:2601.23045

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective

2025-12-19 · Muhammad Osama Imran, Roshni Lulla, Rodney Sappington arxiv

We examine the conceptual and ethical gaps in current representations of Superintelligence misalignment. We find throughout Superintelligence discourse an absent human subject, and an under-developed theorization of an "…

Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce

2026-08-14 · Zeyuan Li, Lukas Petersson, Alessandro Acquisti, Michiel A. Bakker arxiv

Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elic…

Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning

2026-04-04 · Zihong Gao, Hongjian Liang, Lei Hao, Liangjun Ke arxiv

Communication is essential for coordination in \emph{cooperative} multi-agent reinforcement learning under partial observability, yet \emph{cross-timestep} delays cause messages to arrive multiple timesteps after generat…

Multi-agent Reinforcement Learning

Labeling Messages as AI-Generated Does Not Reduce Their Persuasive Effects

2025-04-14 · Isabel O. Gallegos, Chen Shani, Weiyan Shi, Federico Bianchi 외

As generative artificial intelligence (AI) enables the creation and dissemination of information at massive scale and speed, it is increasingly important to understand how people perceive AI-generated content. One promin…

Persuasiveness

Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs

2025-10-13 · Nikita Afonin, Nikita Andriianov, Vahagn Hovhannisyan, Nikhil Bageshpura 외 arxiv

Recent work has shown that narrow finetuning can produce broadly misaligned LLMs, a phenomenon termed emergent misalignment (EM). While concerning, these findings were limited to finetuning and activation steering, leavi…