paper-with-me

Papers

Evaluating calibrated refusal and safe usefulness in dual-use biology settings

2026-07-06 · Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman arxiv

As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk identification and refusal behavior for biological research tasks. The benchmark pairs 61 Routine tasks, legitimate analyses adapted from the published literature, with 46 Red-Team tasks, fictional scenarios that resemble real research but conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates ranged from 7\% to 74\% on Routine tasks and 1\% to 62\% on Red-Team tasks, with many configurations refusing legitimate Routine work at comparable or higher rates than concealed hazards. Refusals were most often triggered by provider API filters applied prior to agentic reasoning. However, models given room to reason showed the potential to identify more real threats. We release BioSecBench-Refusal as a tool for model developers to calibrate capability and caution for agentic biotech R\&D.

📄 PDF Abstract BibTeX arXiv:2607.05462

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization

2024-10-16 · Batuhan K. Karaman, Ishmam Zabir, Alon Benhaim, Vishrav Chaudhary 외

Balancing safety and usefulness in large language models has become a critical challenge in recent years. Models often exhibit unsafe behavior or adopt an overly cautious approach, leading to frequent overrefusal of beni…

Instruction Following

DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models

2025-10-12 · Kaixuan Ren, Preslav Nakov, Usman Naseem arxiv

As vision-language models (VLMs) become increasingly capable, maintaining a balance between safety and usefulness remains a central challenge. Safety mechanisms, while essential, can backfire, causing over-refusal, where…

RAS: Measuring LLM Safety Through Refusal Alignment

2026-06-24 · Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu, Wei-Bin Lee arxiv

Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy. Although useful, output-level evaluat…

Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests

2025-02-08 · David Noever, Forrest McKee

The development of robust safety benchmarks for large language models requires open, reproducible datasets that can measure both appropriate refusal of harmful content and potential over-restriction of legitimate scienti…

valid

Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

2026-05-27 · Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta, Denis Kleyko 외 arxiv

In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations at each transformer block. We find that…