paper-with-me

Papers

Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

2026-05-27 · Richard J. Young, Gregory D. Moody arxiv

A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a working weapon: a keylogger, ransomware, an exploit that runs as written. This asymmetry in the severity of a single act of compliance implies coding-specialized models should clear a higher refusal bar than general-purpose chat models, not a lower one, yet the field cannot tell whether they do. Refusal benchmarks for malicious code are fragmented: they mix requests for executable software with requests for harmful security knowledge and report refusal rates over non-comparable corpora. This paper's central result is that the CODE-versus-KNOWLEDGE classification axis established in a prior four-corpus release remains stable under a substantially expanded corpus pool and an independently refreshed judge panel, evidence that it measures a real construct rather than an artifact of the prompts or judges. Eight corpora spanning diverse elicitation paradigms (direct, jailbreak-decorated, indirect, and agent/interpreter: ASTRA, CySecBench, AdvBench/harmful_behaviors, JailbreakBench, MalwareBench, RedCode, RMCBench, Scam2Prompt) are classified under a five-judge consensus protocol (6,675 prompts x 5 judges = 33,375 calls), reaching Fleiss' kappa = 0.767 [95% CI 0.755, 0.777] ("substantial"). Critically, the panel shares no judge with the prior release (five paid commercial APIs replaced by five open-weight models from five vendors), yet the two panels agree on 94.45% of the 3,133 shared prompts and reach Cohen's kappa = 0.952 [0.942, 0.963] on the 3,031-prompt binary overlap: the axis survives near-total panel replacement. The released bank comprises 4,748 consensus-CODE and 1,923 consensus-KNOWLEDGE prompts, a reliability-quantified benchmark whose central classification axis is shown stable across corpus expansion and judge-panel replacement.

📄 PDF Abstract BibTeX arXiv:2605.28734

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semi-supervised classification by reaching consensus among modalities

2018-05-23 · Zining Zhu, Jekaterina Novikova, Frank Rudzicz

Deep learning has demonstrated abilities to learn complex structures, but they can be restricted by available data. Recently, Consensus Networks (CNs) were proposed to alleviate data sparsity by utilizing features from m…

ClassificationGeneral ClassificationMarketing

Semi-supervised classification by reaching consensus among modalities

2018-10-21 · NIPS Workshop IRASL 2018 · Anonymous

Deep learning has demonstrated abilities to learn complex structures, but they can be restricted by available data. Recently, Consensus Networks (CNs) were proposed to alleviate data sparsity by utilizing features from m…

ClassificationMarketing

Dark Web Activity Classification Using Deep Learning

2023-05-30 · Ali Fayzi, Mohammad Fayzi, Kourosh Dadashtabar Ahmadi

In contemporary times, people rely heavily on the internet and search engines to obtain information, either directly or indirectly. However, the information accessible to users constitutes merely 4% of the overall inform…

ClassificationDeep Learning

Visual Consensus Prompting for Co-Salient Object Detection

2025-04-19 · CVPR 2025 1 · Jie Wang, Nana Yu, Zihao Zhang, Yahong Han

Existing co-salient object detection (CoSOD) methods generally employ a three-stage architecture (i.e., encoding, consensus extraction & dispersion, and prediction) along with a typical full fine-tuning paradigm. Althoug…

Co-Salient Object Detectionobject-detectionObject DetectionSalient Object Detection

VLM-CPL: Consensus Pseudo Labels from Vision-Language Models for Human Annotation-Free Pathological Image Classification

2024-03-23 · Lanfeng Zhong, Xin Liao, Shaoting Zhang, Xiaofan Zhang 외

Despite that deep learning methods have achieved remarkable performance in pathology image classification, they heavily rely on labeled data, demanding extensive human annotation efforts. In this study, we present a nove…

image-classificationImage Classificationzero-shot-classificationZero-Shot Learning