paper-with-me

Papers

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

2026-07-30 · Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu arxiv

Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.

📄 PDF Abstract BibTeX arXiv:2607.27951

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tamper-Resistant Safeguards for Open-Weight LLMs

2024-08-01 · Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou 외

Rapid advances in the capabilities of large language models (LLMs) have raised widespread concerns regarding their potential for malicious use. Open-weight LLMs present unique challenges, as existing safeguards lack robu…

Red TeamingTAR

LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

2024-06-07 · Lukas Helff, Felix Friedrich, Manuel Brack, Kristian Kersting 외

This paper introduces LlavaGuard, a suite of VLM-based vision safeguards that address the critical need for reliable guardrails in the era of large-scale data and models. To this end, we establish a novel open framework,…

Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs

2026-01-20 · Jackson Kaunismaa, Avery Griffin, John Hughes, Christina Q. Knight 외 arxiv

Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be use…

Indiana Jones: There Are Always Some Useful Ancient Relics

2025-01-27 · Junchen Ding, Jiahao Zhang, Yi Liu, Ziqi Ding 외

This paper introduces Indiana Jones, an innovative approach to jailbreaking Large Language Models (LLMs) by leveraging inter-model dialogues and keyword-driven prompts. Through orchestrating interactions among three spec…

Lying Blindly: Bypassing ChatGPT's Safeguards to Generate Hard-to-Detect Disinformation Claims

2024-02-13 · Freddy Heppell, Mehmet E. Bakir, Kalina Bontcheva

As Large Language Models become more proficient, their misuse in coordinated disinformation campaigns is a growing concern. This study explores the capability of ChatGPT with GPT-3.5 to generate short-form disinformation…

World Knowledge