paper-with-me

홈 › Papers

Why Do Language Model Agents Whistleblow?

2025-11-21 · Kushal Agrawal, Frank Xiao, Guido Bergman, Asa Cooper Stickland arxiv

The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds that language models can use tools in ways that contradict the interests or explicit instructions of the user. We study LLM whistleblowing: a subset of this behavior where models disclose suspected misconduct to parties beyond the dialog boundary (e.g., regulatory agencies) without user instruction or knowledge. We introduce an evaluation suite of diverse and realistic staged misconduct scenarios to assess agents for this behavior. Across models and settings, we find that: (1) the frequency of whistleblowing varies widely across model families, (2) increasing the complexity of the task the agent is instructed to complete lowers whistleblowing tendencies, (3) nudging the agent in the system prompt to act morally substantially raises whistleblowing rates, and (4) giving the model more obvious avenues for non-whistleblowing behavior, by providing more tools and a detailed workflow to follow, decreases whistleblowing rates. Additionally, we verify the robustness of our dataset by testing for model evaluation awareness, and find that both black-box methods and probes on model activations show lower evaluation awareness in our settings than in comparable previous work.

📄 PDF Abstract BibTeX arXiv:2511.17085

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Whistleblowing and the machine -- towards a considered position

2026-06-19 · Marija Slavkovik, Liuwen Yu, Leon van der Torre, Reka Markovich arxiv

Artificial intelligent agents and autonomous systems are embedded in our environments. They are both a commercial product and a personal tool that generates a lot of data and can draw conclusions from it: machines genera…

Whistleblower protection in the digital age -- why 'anonymous' is not enough. From technology to a wider view of governance

2021-11-04 · Bettina Berendt, Stefan Schiffner

When technology enters applications and processes with a long tradition of controversial societal debate, multi-faceted new ethical and legal questions arise. This paper focusses on the process of whistleblowing, an acti…

Cultural Vocal Bursts Intensity PredictionFairness

Silencing the Risk, Not the Whistle: A Semi-automated Text Sanitization Tool for Mitigating the Risk of Whistleblower Re-Identification

2024-05-02 · Dimitri Staufer, Frank Pallas, Bettina Berendt

Whistleblowing is essential for ensuring transparency and accountability in both public and private sectors. However, (potential) whistleblowers often fear or face retaliation, even when reporting anonymously. The specif…

Authorship Attribution

A model-based assessment of the cost-benefit balance and the plea bargain in criminality -- A qualitative case study of the Covid-19 epidemic shedding light on the "car wash operation" in Brazil

2022-01-09 · Hyun Mo Yang, Ariana Campos Yang, Silvia Martorano Raimundo

We developed a simple mathematical model to describe criminality and the justice system composed of the police investigation and court trial. The model assessed two features of organized crime -- the cost-benefit analysi…

Mapping Human Anti-collusion Mechanisms to Multi-agent AI Systems

2026-01-01 · Jamiu Idowu, Ahmed Almasoud, Ayman Alfahid arxiv

As multi-agent AI systems become increasingly autonomous, evidence shows they can develop collusive strategies similar to those long observed in human markets and institutions. While human domains have accumulated centur…