US AISI and UK AISI Joint Pre-Deployment Test: Anthropic’s Claude 3.5 Sonnet (October 2024 Release)
This technical report details a pre-deployment evaluation of Anthropic’s upgraded version of Claude 3.5 Sonnet, released October 22, 2024 (hereafter referred to as Sonnet 3.5 (new)). This evaluation was conducted jointly by the United States Artificial Intelligence Safety Institute (US AISI) and the United Kingdom Artificial Intelligence Safety Institute (UK AISI), and this report describes in detail its technical methodology and findings. For general background and a summary of this report, see the corresponding blog post. US AISI and UK AISI’s joint pre-deployment evaluation assessed four domains: biological capabilities, cyber capabilities, software and AI development capabilities, and safeguard effectiveness. US AISI and UK AISI each ran independent tests on Sonnet 3.5 (new), working together to inform and improve methodology and interpretation of findings. US AISI and UK AISI shared their initial findings with Anthropic prior to the model’s release. The following sections introduce each evaluation domain jointly and present specific technical descriptions, methodologies, and findings in each domain as specific to either US AISI or UK AISI, as appropriate.
Code (0)
등록된 구현이 없습니다.
Tasks
Similar Papers 제목 키워드 기반
The lexical and grammatical sources of neg-raising inferences
We investigate neg(ation)-raising inferences, wherein negation on a predicate can be interpreted as though in that predicate's subordinate clause. To do this, we collect a large-scale dataset of neg-raising judgments for…
NegationEstimating Early Fundraising Performance of Innovations via Graph-based Market Environment Model
Well begun is half done. In the crowdfunding market, the early fundraising performance of the project is a concerned issue for both creators and platforms. However, estimating the early fundraising performance before the…
NodMAISI: Nodule-Oriented Medical AI for Synthetic Imaging
Objective: Although medical imaging datasets are increasingly available, abnormal and annotation-intensive findings critical to lung cancer screening, particularly small pulmonary nodules, remain underrepresented and inc…
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This paper introduces Transient Turn Injection(TTI), a new multi-turn attack te…
Adversarial RobustnessA Comparison of Praising Skills in Face-to-Face and Remote Dialogues
Praising behavior is considered to an important method of communication in daily life and social activities. An engineering analysis of praising behavior is therefore valuable. However, a dialogue corpus for this analysi…