paper-with-me

홈 › Papers

The Safeguard Worked. Is the LLM System Safer?

2026-09-01 · Pingyu Wu, Weiming Zhang, Nenghai Yu hf

Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deployment criterion. The evidence requirements are strongly asymmetric. One attack that obtains harmful help from the deployed service suffices to establish that such help remains, and such attacks appear repeatedly in the coded record. Establishing that little remains cannot follow from the safeguard's own numbers alone; it also requires evidence about what the surrounding system still allows after the safeguard performs its local function. Such evidence is supported or derived in only a small minority of the depth-coded claims, and one such claim bounds its scoped residual. A better local score is therefore not, by itself, a stronger claim about the deployment. Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer.

📄 PDF Abstract BibTeX arXiv:2609.00519

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Libra: Large Chinese-based Safeguard for AI Content

2025-07-29 · Ziyang Chen, Huimu Yu, Xing Wu, Dongqin Liu 외 arxiv

Large language models (LLMs) excel in text understanding and generation but raise significant safety and ethical concerns in high-stakes applications. To mitigate these risks, we present Libra-Guard, a cutting-edge safeg…

Driving-Policy Adaptive Safeguard for Autonomous Vehicles Using Reinforcement Learning

2020-12-02 · Zhong Cao, Shaobing Xu, Songan Zhang, Huei Peng 외

Safeguard functions such as those provided by advanced emergency braking (AEB) can provide another layer of safety for autonomous vehicles (AV). A smart safeguard function should adapt the activation conditions to the dr…

Autonomous VehiclesCollision Avoidancereinforcement-learningReinforcement Learning+1

Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

2023-08-25 · Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov 외

With the rapid evolution of large language models (LLMs), new and hard-to-predict harmful capabilities are emerging. This requires developers to be able to identify risks through the evaluation of "dangerous capabilities…

SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models

2025-10-08 · Huahui Yi, Kun Wang, Qiankun Li, Miao Yu 외 arxiv

Multimodal Large Reasoning Models (MLRMs) demonstrate impressive cross-modal reasoning but often amplify safety risks under adversarial or unsafe prompts, a phenomenon we call the \textit{Reasoning Tax}. Existing defense…

Reinforcement LearningMultimodal Reasoning

Learning-Augmented Decentralized Online Convex Optimization in Networks

2023-06-16 · Pengfei Li, Jianyi Yang, Adam Wierman, Shaolei Ren

This paper studies decentralized online convex optimization in a networked multi-agent system and proposes a novel algorithm, Learning-Augmented Decentralized Online optimization (LADO), for individual agents to select a…