paper-with-me

홈 › Papers

HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment

2024-11-11 · Yannis Belkhiter, Giulio Zizzo, Sergio Maffeis

With the introduction of the transformers architecture, LLMs have revolutionized the NLP field with ever more powerful models. Nevertheless, their development came up with several challenges. The exponential growth in computational power and reasoning capabilities of language models has heightened concerns about their security. As models become more powerful, ensuring their safety has become a crucial focus in research. This paper aims to address gaps in the current literature on jailbreaking techniques and the evaluation of LLM vulnerabilities. Our contributions include the creation of a novel dataset designed to assess the harmfulness of model outputs across multiple harm levels, as well as a focus on fine-grained harm-level analysis. Using this framework, we provide a comprehensive benchmark of state-of-the-art jailbreaking attacks, specifically targeting the Vicuna 13B v1.5 model. Additionally, we examine how quantization techniques, such as AWQ and GPTQ, influence the alignment and robustness of models, revealing trade-offs between enhanced robustness with regards to transfer attacks and potential increases in vulnerability on direct ones. This study aims to demonstrate the influence of harmful input queries on the complexity of jailbreaking techniques, as well as to deepen our understanding of LLM vulnerabilities and improve methods for assessing model robustness when confronted with harmful content, particularly in the context of compression strategies.

📄 PDF Abstract BibTeX arXiv:2411.06835

Code (0)

등록된 구현이 없습니다.

Tasks

Quantization

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences

2025-05-30 · Mingqian Zheng, Wenjia Hu, Patrick Zhao, Motahhare Eslami 외

Current LLMs are trained to refuse potentially harmful input queries regardless of whether users actually had harmful intents, causing a tradeoff between safety and user experience. Through a study of 480 participants ev…

An Open Knowledge Graph-Based Approach for Mapping Concepts and Requirements between the EU AI Act and International Standards

2024-08-21 · Julio Hernandez, Delaram Golpayegani, Dave Lewis

The many initiatives on trustworthy AI result in a confusing and multipolar landscape that organizations operating within the fluid and complex international value chains must navigate in pursuing trustworthy AI. The EU'…

Knowledge GraphsNavigate

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

2026-05-01 · Yutao Hou, Yihan Jiang, Yuhan Xie, Jian Yang 외 arxiv

Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal activities or unethical behavior, posing serious compliance risks. T…

TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors

2026-09-04 · Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho arxiv

Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful prompts with varying threat implicitness. We introduce TIER, a Threat Implicitness Benchmark for behavioral safety e…

InvisibleBench: A Deployment Gate for Caregiving Relationship AI

2025-11-25 · Ali Madad arxiv

InvisibleBench is a deployment gate for caregiving-relationship AI, evaluating 3-20+ turn interactions across five dimensions: Safety, Compliance, Trauma-Informed Design, Belonging/Cultural Fitness, and Memory. The bench…