paper-with-me

홈 › Papers

UK AISI Alignment Evaluation Case-Study

2026-04-01 · Alexandra Souly, Robert Kirk, Jacob Merizian, Abby D'Cruz, Xander Davies arxiv

This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate whether frontier models sabotage safety research when deployed as coding assistants within an AI lab. Applying our methods to four frontier models, we find no confirmed instances of research sabotage. However, we observe that Claude Opus 4.5 Preview (a pre-release snapshot of Opus 4.5) and Sonnet 4.5 frequently refuse to engage with safety-relevant research tasks, citing concerns about research direction, involvement in self-training, and research scope. We additionally find that Opus 4.5 Preview shows reduced unprompted evaluation awareness compared to Sonnet 4.5, while both models can distinguish evaluation from deployment scenarios when prompted. Our evaluation framework builds on Petri, an open-source LLM auditing tool, with a custom scaffold designed to simulate realistic internal deployment of a coding agent. We validate that this scaffold produces trajectories that all tested models fail to reliably distinguish from real deployment data. We test models across scenarios varying in research motivation, activity type, replacement threat, and model autonomy. Finally, we discuss limitations including scenario coverage and evaluation awareness.

📄 PDF Abstract BibTeX arXiv:2604.00788

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CVC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models

2025-06-02 · Ping Wu, Guobin Shen, Dongcheng Zhao, Yuwei Wang 외

Ensuring that Large Language Models (LLMs) align with mainstream human values and ethical norms is crucial for the safe and sustainable development of AI. Current value evaluation and alignment are constrained by Western…

Benchmarking

Raising the TM Threshold in Neural MT Post-Editing: a Case Study onTwo Datasets

2019-08-01 · WS 2019 8 · Anna Zaretskaya

Zoom In Disparities in Healthcare LLM Q&A

2025-10-20 · Ipek Baris Schlicht, Burcu Sayin, Zhixue Zhao, Frederik M. Labonté 외 arxiv

Equitable access to reliable health information is vital when integrating AI into healthcare. Yet, information quality varies across languages, raising concerns about the reliability and consistency of multilingual Large…

US AISI and UK AISI Joint Pre-Deployment Test: Anthropic’s Claude 3.5 Sonnet (October 2024 Release)

2024-11-19 · NIST 2024 11 · US AI Safety Institute, UK AI Safety Institute

This technical report details a pre-deployment evaluation of Anthropic’s upgraded version of Claude 3.5 Sonnet, released October 22, 2024 (hereafter referred to as Sonnet 3.5 (new)). This evaluation was conducted jointly…

A Revealed Preference Framework for AI Alignment

2026-03-29 · Elchin Suleymanov arxiv

Human decision makers increasingly delegate choices to AI agents, raising a natural question: does the AI implement the human principal's preferences or pursue its own? To study this question using revealed preference te…