paper-with-me

홈 › Papers

Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation

2024-07-11 · Riccardo Cantini, Giada Cosenza, Alessio Orsino, Domenico Talia

Large Language Models (LLMs) have revolutionized artificial intelligence, demonstrating remarkable computational power and linguistic capabilities. However, these models are inherently prone to various biases stemming from their training data. These include selection, linguistic, and confirmation biases, along with common stereotypes related to gender, ethnicity, sexual orientation, religion, socioeconomic status, disability, and age. This study explores the presence of these biases within the responses given by the most recent LLMs, analyzing the impact on their fairness and reliability. We also investigate how known prompt engineering techniques can be exploited to effectively reveal hidden biases of LLMs, testing their adversarial robustness against jailbreak prompts specially crafted for bias elicitation. Extensive experiments are conducted using the most widespread LLMs at different scales, confirming that LLMs can still be manipulated to produce biased or inappropriate responses, despite their advanced capabilities and sophisticated alignment processes. Our findings underscore the importance of enhancing mitigation techniques to address these safety issues, toward a more sustainable and inclusive artificial intelligence.

📄 PDF Abstract BibTeX arXiv:2407.08441

Code (1)

SCAlabUnical/LLM-Bias-Jailbreak 공식 구현

Tasks

Adversarial RobustnessFairnessPrompt Engineering

Similar Papers 제목 키워드 기반

Is the System Message Really Important to Jailbreaks in Large Language Models?

2024-02-20 · Xiaotian Zou, Yongkang Chen, Ke Li

The rapid evolution of Large Language Models (LLMs) has rendered them indispensable in modern society. While security measures are typically to align LLMs with human values prior to release, recent studies have unveiled …

Evolutionary Algorithms

BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models

2024-10-17 · Isack Lee, Haebin Seong

Although large language models (LLMs) demonstrate impressive proficiency in various tasks, they present potential safety risks, such as `jailbreaks', where malicious inputs can coerce LLMs into generating harmful content…

Red TeamingSafety AlignmentText Generation

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models

2025-01-30 · Haoyu Liang, Youran Sun, Yunfeng Cai, Jun Zhu 외

The security issue of large language models (LLMs) has gained wide attention recently, with various defense mechanisms developed to prevent harmful output, among which safeguards based on text embedding models serve as a…

LLM Jailbreak Detection for (Almost) Free!

2025-09-18 · Guorui Chen, Yifan Xia, Xiaojun Jia, Zhijiang Li 외 arxiv

Large language models (LLMs) enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak detection methods show promise in mitiga…

Quantitative Certification of Bias in Large Language Models

2024-05-29 · Isha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi 외

Large Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly evaluate LLM bias, as they can not scale to large number of in…

Benchmarking