paper-with-me

홈 › Papers

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

2026-07-08 · Lifei Liu, Haoran Yu, Xiaochong Jiang, Su Wang, Pin Qian, Yihang Chen arxiv

Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may be reframed as plausible operational work, the planner may refuse or transform the request, and the executor may act under delegation prompts implying prior approval. To separate these factors, we introduce a five-condition controlled contrast design, evaluated on 30 synthetic harmful scenarios and an exploratory external validation set from four agent-safety benchmarks using LLM-judged compliance. Our results show that aggregate pipeline safety is not a stable architectural property. Operational reframing is the most portable risk signal, increasing compliance for GPT, Gemini, and DeepSeek across both scenario sets, while Claude is comparatively resistant. Planner behavior can offset this risk mainly through refusal; however, when the planner produces executable steps, the executor may become more compliant than under the direct operational baseline. Approval-framed delegation is sensitive to prompt design, model pairing, and scenario source, and a skeptical executor prompt sharply reduces compliance. Raw-direct model rankings can also mispredict deployed planner-executor behavior. Gemini is safest under raw direct prompts in the primary set yet shows the largest amplification with a Claude planner, rising from 8.9 percent to 38.9 percent compliance. GPTs near-zero aggregate pipeline effect instead hides a reframing increase canceled by planner refusal. These findings suggest that multi-agent safety evaluations should report reframing, planner behavior, delegation framing, and model pairing separately before attributing failures to architecture itself.

📄 PDF Abstract BibTeX arXiv:2607.07097

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet Named

2026-04-21 · Rohith Nama arxiv

Autonomous AI agents now plan, decide, and act on behalf of users across healthcare, financial services, and workplace contexts, often without step-by-step human approval. Existing AI literacy frameworks were built for a…

Controlled Neural Sentence-Level Reframing of News Articles

2021-09-10 · Findings (EMNLP) 2021 11 · Wei-Fan Chen, Khalid Al-Khatib, Benno Stein, Henning Wachsmuth

Framing a news article means to portray the reported event from a specific perspective, e.g., from an economic or a health perspective. Reframing means to change this perspective. Depending on the audience or the submess…

ArticlesSentenceText Generation

Delegation in Veto Bargaining

2020-06-11 · Navin Kartik, Andreas Kleiner, Richard Van Weelden

A proposer requires the approval of a veto player to change a status quo. Preferences are single peaked. Proposer is uncertain about Vetoer's ideal point. We study Proposer's optimal mechanism without transfers. Vetoer i…

Reframing Instructional Prompts to GPTk's Language

2021-09-16 · Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi 외

What kinds of instructional prompts are easier to follow for Language Models (LMs)? We study this question by conducting extensive empirical analysis that shed light on important features of successful instructional prom…

Few-Shot LearningQuestion GenerationZero-Shot Learning

Reframing Instructional Prompts to GPTk’s Language

2022-05-01 · Findings (ACL) 2022 5 · Daniel Khashabi, Chitta Baral, Yejin Choi, Hannaneh Hajishirzi

What kinds of instructional prompts are easier to follow for Language Models (LMs)? We study this question by conducting extensive empirical analysis that shed light on important features of successful instructional prom…