paper-with-me

홈 › Papers

Great Models Think Alike and this Undermines AI Oversight

2025-02-06 · Shashwat Goel, Joschka Struber, Ilze Amanda Auzina, Karuna K Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, Jonas Geiping

As Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans. There is hope that other language models can automate both these tasks, which we refer to as "AI Oversight". We study how model similarity affects both aspects of AI oversight by proposing a probabilistic metric for LM similarity based on overlap in model mistakes. Using this metric, we first show that LLM-as-a-judge scores favor models similar to the judge, generalizing recent self-preference results. Then, we study training on LM annotations, and find complementary knowledge between the weak supervisor and strong student model plays a crucial role in gains from "weak-to-strong generalization". As model capabilities increase, it becomes harder to find their mistakes, and we might defer more to AI oversight. However, we observe a concerning trend -- model mistakes are becoming more similar with increasing capabilities, pointing to risks from correlated failures. Our work underscores the importance of reporting and correcting for model similarity, especially in the emerging paradigm of AI oversight.

📄 PDF Abstract BibTeX arXiv:2502.04313

Code (2)

model-similarity/lm-similarity 공식 구현 pytorch
model-similarity/lm-similarity/tree/main/applications 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Accelerating Hybrid Agent-Based Models and Fuzzy Cognitive Maps: How to Combine Agents who Think Alike?

2024-09-01 · Philippe J. Giabbanelli, Jack T. Beerman

While Agent-Based Models can create detailed artificial societies based on individual differences and local context, they can be computationally intensive. Modelers may offset these costs through a parsimonious use of th…

Community DetectionGPU

Accountability Capture: How Record-Keeping to Support AI Transparency and Accountability (Re)shapes Algorithmic Oversight

2025-10-06 · Shreya Chappidi, Jennifer Cobbe, Chris Norval, Anjali Mazumder 외 arxiv

Accountability regimes typically encourage record-keeping to enable the transparency that supports oversight, investigation, contestation, and redress. However, implementing such record-keeping can introduce consideratio…

Do Deep Minds Think Alike? Selective Adversarial Attacks for Fine-Grained Manipulation of Multiple Deep Neural Networks

2020-03-26 · Zain Khan, Jirong Yi, Raghu Mudumbai, Xiaodong Wu 외

Recent works have demonstrated the existence of {\it adversarial examples} targeting a single machine learning system. In this paper we ask a simple but fundamental question of "selective fooling": given {\it multiple} m…

BIG-bench Machine Learning

The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting

2026-05-08 · Lauri Lovén, Sasu Tarkoma arxiv

Eliciting truthful reports from autonomous agents is a core problem in scalable AI oversight: a principal scores the agent's report using a strictly proper scoring rule, but the agent also benefits from the report throug…

From Democracies to Autocracies: How AI Systems Enable Authoritarianism by Design

2026-06-15 · Jeba Sania, Marta Ziosi, Fazl Barez arxiv

AI-enabled authoritarianism is not confined to autocracies. In this paper, we provide greater transparency by investigating and mapping the lifecycles of six AI systems deployed in different political regimes, ranging fr…