paper-with-me

Papers

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture

2025-12-09 · Gary Ackerman, Brandon Behlendorf, Zachary Kallenborn, Sheriff Almakki, Doug Clifford, Jenna LaTourette, Hayley Peterson, Noah Sheinbaum, Olivia Shoemaker, Anna Wetzel arxiv

Both model developers and policymakers seek to quantify and mitigate the risk of rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons. An important element of such efforts is the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper describes the first component of a novel Biothreat Benchmark Generation (BBG) Framework. The BBG approach is designed to help model developers and evaluators reliably measure and assess the biosecurity risk uplift and general harm potential of existing and future AI models, while accounting for key aspects of the threat itself that are often overlooked in other benchmarking efforts, including different actor capability levels, and operational (in addition to purely technical) risk factors. As a pilot, the BBG is first being developed to address bacterial biological threats only. The BBG is built upon a hierarchical structure of biothreat categories, elements and tasks, which then serves as the basis for the development of task-aligned queries. This paper outlines the development of this biothreat task-query architecture, which we have named the Bacterial Biothreat Schema, while future papers will describe follow-on efforts to turn queries into model prompts, as well as how the resulting benchmarks can be implemented for model evaluation. Overall, the BBG Framework, including the Bacterial Biothreat Schema, seeks to offer a robust, re-usable structure for evaluating bacterial biological risks arising from LLMs across multiple levels of aggregation, which captures the full scope of technical and operational requirements for biological adversaries, and which accounts for a wide spectrum of biological adversary capabilities.

📄 PDF Abstract BibTeX arXiv:2512.08130

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) Dataset

2025-12-09 · Gary Ackerman, Theodore Wilson, Zachary Kallenborn, Olivia Shoemaker 외 arxiv

The potential for rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons has generated significant policy, acade…

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models II: Benchmark Generation Process

2025-12-09 · Gary Ackerman, Zachary Kallenborn, Anna Wetzel, Hayley Peterson 외 arxiv

The potential for rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons has generated significant policy, acade…

Red Teaming

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

2026-07-21 · Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon 외 hf

Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims we…

ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry

2025-07-22 · Tianze Xu, Pengrui Lu, Lyumanshan Ye, Xiangkun Hu 외 arxiv

The emergence of deep research systems presents significant capabilities in problem-solving, extending from basic queries to sophisticated research tasks. However, existing benchmarks primarily evaluate these systems as …

Insights from Benchmarking Frontier Language Models on Web App Code Generation

2024-09-08 · Yi Cui

This paper presents insights from evaluating 16 frontier large language models (LLMs) on the WebApp1K benchmark, a test suite designed to assess the ability of LLMs to generate web application code. The results reveal th…

BenchmarkingCode GenerationPrompt Engineering