paper-with-me

홈 › Papers

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

2026-07-20 · Zhida He, Xia Hu, Baichen Le, Chunxiao Li, Jiajia Li, Lijun Li, Chaochao Lu, Jing Shao, Youbang Sun, Hua Tang, Xiang Wang, Xiao Wang, Xiaoyu Wen, Tong Wu, Jia Xu, Peng Yu, Shu Yu, Jie Zhang, Qiaosheng Zhang, Yi Zhang, Xing-Ming Zhao, Tianhang Zheng, Ziyuan Zhou arxiv

Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace current safeguards. To assess the biological risks of frontier models, we develop Intern-BioBreaker, a specialized bio-red-teaming model, together with an integrated computational-to-physical framework that couples model-level stress testing with wet-lab validation. Within this framework, Intern-BioBreaker generates targeted jailbreak prompts to test whether aligned models can be induced to provide operational guidance for safety-sensitive biological tasks or produce sequence-level outputs with potentially harmful properties. Selected sequence outputs are then carried forward for DNA synthesis, host expression, and orthogonal protein verification to assess whether model-generated designs can yield the intended biological products. Our evaluation reveals a concerning gap between text-level safeguards and the risks posed by capable scientific models: (i) Intern-BioBreaker outperforms baseline attack models and reveals widespread bio-risk jailbreak vulnerabilities across both open-weight and proprietary frontier LLMs, with several targets reaching near-saturated or 100% task-level attack success rate (ASR); (ii) in sequence-level case studies, GPT-5.5 can be induced to generate modified viral candidate sequences with pathogenic potential; the corresponding translated proteins may exhibit even stronger receptor-binding affinity and thus enhanced infection potential; and (iii) end-to-end verification shows that selected model-generated biological designs are not merely textual artifacts, but can be physically realized under controlled experimental settings. These findings underscore the need for stronger biological red-teaming, nucleic acid synthesis screening, and safety mechanisms that keep pace with model capabilities.

📄 PDF Abstract BibTeX arXiv:2607.18056

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measuring Biological Capabilities and Risks of AI Agents

2026-06-18 · Patricia Paskov, Jeffrey Lee, Kyle Brady, Alyssa Worland arxiv

This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or col…

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

2025-07-22 · Shanghai AI Lab, :, Xiaoyang Chen, Yunhao Chen 외 arxiv

To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, this report presents a comprehensive assessment of their frontier risks. Drawing on the E-T-C analysis (d…

Evaluating Frontier Models for Dangerous Capabilities

2024-03-20 · Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan 외

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 mode…

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

2026-06-09 · Andrew Bo Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman 외 arxiv

Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM agents can also perform in silico biolo…

Weather Emulators at the Frontier of Heat Extremes Predictability

2026-07-30 · Cas Decancq, Thomas Mortier, Jessica Keune, Diego G. Miralles arxiv

Atmospheric predictability declines rapidly beyond the next ten days, such that forecasts at longer lead times primarily convey large-scale trends rather than specific states. Yet in a warming world, improving early warn…