paper-with-me

홈 › Papers

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

2025-07-04 · Christopher Summerfield, Lennart Luettgau, Magda Dubois, Hannah Rose Kirk, Kobi Hackenburg, Catherine Fist, Katarina Slama, Nicola Ding, Rebecca Anselmetti, Andrew Strait, Mario Giulianelli, Cozmin Ududec arxiv

We examine recent research that asks whether current AI systems may be developing a capacity for "scheming" (covertly and strategically pursuing misaligned goals). We compare current research practices in this field to those adopted in the 1970s to test whether non-human primates could master natural language. We argue that there are lessons to be learned from that historical research endeavour, which was characterised by an overattribution of human traits to other agents, an excessive reliance on anecdote and descriptive analysis, and a failure to articulate a strong theoretical framework for the research. We recommend that research into AI scheming actively seeks to avoid these pitfalls. We outline some concrete steps that can be taken for this research programme to advance in a productive and scientifically rigorous fashion.

📄 PDF Abstract BibTeX arXiv:2507.03409

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Frontier Models are Capable of In-context Scheming

2024-12-06 · Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni 외

Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives - also known as schemi…

Evaluating and Understanding Scheming Propensity in LLM Agents

2026-03-02 · Mia Hopman, Jannes Elstner, Maria Avramidou, Amritanshu Prasad 외 arxiv

As frontier language models are increasingly deployed as autonomous agents pursuing complex, long-term objectives, there is increased risk of scheming: agents covertly pursuing misaligned goals. Prior work has focused on…

Scheming Ability in LLM-to-LLM Strategic Interactions

2025-10-11 · Thao Pham arxiv

As large language model (LLM) agents are deployed autonomously in diverse contexts, evaluating their capacity for strategic deception becomes crucial. While recent research has examined how AI systems scheme against huma…

Learning to Fuse Temporal Proximity Networks: A Case Study in Chimpanzee Social Interactions

2025-02-01 · Yixuan He, Aaron Sandel, David Wipf, Mihai Cucuringu 외

How can we identify groups of primate individuals which could be conjectured to drive social structure? To address this question, one of us has collected a time series of data for social interactions between chimpanzees.…

Time Series

Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence

2026-04-10 · Tommy Shaffer Shane, Simon Mylius, Hamish Hobbs arxiv

Scheming, the covert pursuit of misaligned goals by AI systems, represents a potentially catastrophic risk, yet scheming research suffers from significant limitations. In particular, scheming evaluations demonstrate beha…