paper-with-me

Papers

Misalignment Bounty: Crowdsourcing AI Agent Misbehavior

2025-10-22 · Rustem Turtayev, Natalia Fedorova, Oleg Serikov, Sergey Koldyba, Lev Avagyan, Dmitrii Volkov arxiv

Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents pursuing unintended or unsafe goals. The bounty received 295 submissions, of which nine were awarded. This report explains the program's motivation and evaluation criteria, and walks through the nine winning submissions step by step.

📄 PDF Abstract BibTeX arXiv:2510.19738

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Training Agents to Self-Report Misbehavior

2026-02-25 · Bruce W. Lee, Chen Yueh-Han, Tomek Korbak arxiv

Frontier AI agents may pursue hidden goals while concealing their pursuit from oversight. Alignment training aims to prevent such behavior by reinforcing the correct goals, but alignment may not always succeed and can le…

Artificial Bugs for Crowdsearch

2024-03-14 · Hans Gersbach, Fikri Pitsuwan, Pio Blieske

Bug bounty programs, where external agents are invited to search and report vulnerabilities (bugs) in exchange for rewards (bounty), have become a major tool for companies to improve their systems. We suggest augmenting …

Gram: Assessing sabotage propensities via automated alignment auditing

2026-05-28 · David Lindner, Victoria Krakovna, Sebastian Farquhar arxiv

We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models across 17 simulated agentic deployment scenarios that incentivize sabota…

Wink: Recovering from Misbehaviors in Coding Agents

2026-02-19 · Rahul Nanda, Chandra Maddila, Smriti Jha, Euna Mehnaz Khan 외 arxiv

Autonomous coding agents, powered by large language models (LLMs), are increasingly being adopted in the software industry to automate complex engineering tasks. However, these agents are prone to a wide range of misbeha…

Agent Hunt: Bounty Based Collaborative Autoformalization With LLM Agents

2026-03-06 · Chad E. Brown, Cezary Kaliszyk, Josef Urban arxiv

We describe an experiment in large-scale autoformalization of algebraic topology in an Interactive Theorem Proving (ITP) environment, where the workload is distributed among multiple LLM-based coding agents. Rather than …