paper-with-me

홈 › Papers

Mitigating Cyber Risk in the Age of Open-Weight LLMs: Policy Gaps and Technical Realities

2025-05-21 · Alfonso de Gregorio

Open-weight general-purpose AI (GPAI) models offer significant benefits but also introduce substantial cybersecurity risks, as demonstrated by the offensive capabilities of models like DeepSeek-R1 in evaluations such as MITRE's OCCULT. These publicly available models empower a wider range of actors to automate and scale cyberattacks, challenging traditional defence paradigms and regulatory approaches. This paper analyzes the specific threats -- including accelerated malware development and enhanced social engineering -- magnified by open-weight AI release. We critically assess current regulations, notably the EU AI Act and the GPAI Code of Practice, identifying significant gaps stemming from the loss of control inherent in open distribution, which renders many standard security mitigations ineffective. We propose a path forward focusing on evaluating and controlling specific high-risk capabilities rather than entire models, advocating for pragmatic policy interpretations for open-weight systems, promoting defensive AI innovation, and fostering international collaboration on standards and cyber threat intelligence (CTI) sharing to ensure security without unduly stifling open technological progress.

📄 PDF Abstract BibTeX arXiv:2505.17109

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

2025-08-05 · Eric Wallace, Olivia Watkins, Miles Wang, Kai Chen 외 arxiv

In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in…

Ollabench: Evaluating LLMs' Reasoning for Human-centric Interdependent Cybersecurity

2024-06-11 · Tam N. Nguyen

Large Language Models (LLMs) have the potential to enhance Agent-Based Modeling by better representing complex interdependent cybersecurity systems, improving cybersecurity threat modeling and risk management. However, e…

Aggressive Compression Enables LLM Weight Theft

2026-01-03 · Davis Brown, Juan-Pablo Rivera, Dan Hendrycks, Mantas Mazeika arxiv

As frontier AIs become more powerful and costly to develop, adversaries have increasing incentives to steal model weights by mounting exfiltration attacks. In this work, we consider exfiltration attacks where an adversar…

An Independent Safety Evaluation of Kimi K2.5

2026-04-03 · Zheng-Xin Yong, Parv Mahajan, Andy Wang, Ida Caspary 외 arxiv

Kimi K2.5 is an open-weight LLM that rivals closed models across coding, multimodal, and agentic benchmarks, but was released without an accompanying safety evaluation. In this work, we conduct a preliminary safety asses…

Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

2023-12-07 · Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan 외

This paper presents CyberSecEval, a comprehensive benchmark developed to help bolster the cybersecurity of Large Language Models (LLMs) employed as coding assistants. As what we believe to be the most extensive unified c…

Language ModelingLanguage ModellingLarge Language Model