paper-with-me

홈 › Papers

SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation

2025-08-08 · Lai Jiang, Yuekang Li, Xiaohan Zhang, Youtao Ding, Li Pan arxiv

Accurate jailbreak evaluation is critical for LLM red team testing and jailbreak research. Mainstream methods rely on binary classification (string matching, toxic text classifiers, and LLM-based methods), outputting only "yes/no" labels without quantifying harm severity. Emerged multi-dimensional frameworks (e.g., Security Violation, Relative Truthfulness and Informativeness) use unified evaluation standards across scenarios, leading to scenario-specific mismatches (e.g., "Relative Truthfulness" is irrelevant to "hate speech"), undermining evaluation accuracy. To address these, we propose SceneJailEval, with key contributions: (1) A pioneering scenario-adaptive multi-dimensional framework for jailbreak evaluation, overcoming the critical "one-size-fits-all" limitation of existing multi-dimensional methods, and boasting robust extensibility to seamlessly adapt to customized or emerging scenarios. (2) A novel 14-scenario dataset featuring rich jailbreak variants and regional cases, addressing the long-standing gap in high-quality, comprehensive benchmarks for scenario-adaptive evaluation. (3) SceneJailEval delivers state-of-the-art performance with an F1 score of 0.917 on our full-scenario dataset (+6% over SOTA) and 0.995 on JBB (+3% over SOTA), breaking through the accuracy bottleneck of existing evaluation methods in heterogeneous scenarios and solidifying its superiority.

📄 PDF Abstract BibTeX arXiv:2508.06194

Code (0)

등록된 구현이 없습니다.

Tasks

Binary Classification

Similar Papers 제목 키워드 기반

Adaptive Testing for Connected and Automated Vehicles with Sparse Control Variates in Overtaking Scenarios

2022-07-19 · Jingxuan Yang, Honglin He, Yi Zhang, Shuo Feng 외

Testing and evaluation is a critical step in the development and deployment of connected and automated vehicles (CAVs). Due to the black-box property and various types of CAVs, how to test and evaluate CAVs adaptively re…

regression

Croppable Knowledge Graph Embedding

2024-07-03 · Yushan Zhu, Wen Zhang, Zhiqiang Liu, Mingyang Chen 외

Knowledge Graph Embedding (KGE) is a common method for Knowledge Graphs (KGs) to serve various artificial intelligence tasks. The suitable dimensions of the embeddings depend on the storage and computing conditions of th…

Graph EmbeddingKnowledge Graph EmbeddingKnowledge GraphsLanguage Modelling

Adaptive Safety Evaluation for Connected and Automated Vehicles with Sparse Control Variates

2022-12-01 · Jingxuan Yang, Haowei Sun, Honglin He, Yi Zhang 외

Safety performance evaluation is critical for developing and deploying connected and automated vehicles (CAVs). One prevailing way is to design testing scenarios using prior knowledge of CAVs, test CAVs in these scenario…

DiCriTest: Testing Scenario Generation for Decision-Making Agents Considering Diversity and Criticality

2025-08-15 · Qitong Chu, Yufeng Yue, Danya Yao, Huaxin Pei arxiv

The growing deployment of decision-making agents in dynamic environments increases the demand for safety verification. While critical testing scenario generation has emerged as an appealing verification methodology, effe…

Dimensionality Reduction

Cardinality Estimation for High Dimensional Similarity Queries with Adaptive Bucket Probing

2026-04-06 · Zhonghan Chen, Qintian Guo, Ruiyuan Zhang, Xiaofang Zhou arxiv

In this work, we address the problem of cardinality estimation for similarity search in high-dimensional spaces. Our goal is to design a framework that is lightweight, easy to construct, and capable of providing accurate…