paper-with-me

홈 › Papers

From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents

2026-02-04 · Xinyue Wang, Yuanhe Zhang, Zhengshuo Gong, Haoran Gao, Fanyu Meng, Zhenhong Zhou, Li Sun, Yang Liu, Sen Su arxiv

The enhanced capabilities of LLM-based agents come with an emergency for model planning and tool-use abilities. Attributing to helpful-harmless trade-off from LLM alignment, agents typically also inherit the flaw of "over-refusal", which is a passive failure mode. However, the proactive planning and action capabilities of agents introduce another crucial danger on the other side of the trade-off. This phenomenon we term "Toxic Proactivity'': an active failure mode in which an agent, driven by the optimization for Machiavellian helpfulness, disregards ethical constraints to maximize utility. Unlike over-refusal, Toxic Proactivity manifests as the agent taking excessive or manipulative measures to ensure its "usefulness'' is maintained. Existing research pays little attention to identifying this behavior, as it often lacks the subtle context required for such strategies to unfold. To reveal this risk, we introduce a novel evaluation framework based on dilemma-driven interactions between dual models, enabling the simulation and analysis of agent behavior over multi-step behavioral trajectories. Through extensive experiments with mainstream LLMs, we demonstrate that Toxic Proactivity is a widespread behavioral phenomenon and reveal two major tendencies. We further present a systematic benchmark for evaluating Toxic Proactive behavior across contextual settings.

📄 PDF Abstract BibTeX arXiv:2602.04197

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment

2026-02-16 · Laurène Vaugrante, Anietta Weckauff, Thilo Hagendorff arxiv

Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misalignment". Moreover, research has shown t…

Knowing Isn't Understanding: Re-grounding Generative Proactivity with Epistemic and Behavioral Insight

2026-02-16 · Kirandeep Kaur, Xingda Lyu, Chirag Shah arxiv

Generative AI agents equate understanding with resolving explicit queries, an assumption that confines interaction to what users can articulate. This assumption breaks down when users themselves lack awareness of what is…

BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum

2025-05-27 · Yubin Kim, Zhiyuan Hu, Hyewon Jeong, Eugene Park 외

Large Language Models (LLMs) as clinical agents require careful behavioral adaptation. While adept at reactive tasks (e.g., diagnosis reasoning), LLMs often struggle with proactive engagement, like unprompted identificat…

LCAM: A Framework for Diagnosing Interactional Alignment Failures in Con-versational AI

2026-06-06 · Manuele Reani, Hongyu Tian arxiv

Conversational AI is increasingly used for advice, interpretation, reassurance, and decision support in contexts where users may be vulnerable, uncertain, or dependent on the system's apparent competence. Existing alignm…

Persona-Model Collapse in Emergent Misalignment

2026-05-13 · Davi Bastos Costa, Renato Vicente arxiv

Fine-tuning large language models on narrow data with harmful content produces broadly misaligned behavior on unrelated prompts, a phenomenon known as emergent misalignment. We propose that emergent misalignment involves…