paper-with-me

홈 › Papers

Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

2026-03-11 · Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali arxiv

The rapid adoption of large language models (LLMs) in financial services introduces new operational, regulatory, and security risks. Yet most red-teaming benchmarks remain domain-agnostic and fail to capture failure modes specific to regulated BFSI settings, where harmful behavior can be elicited through legally or professionally plausible framing. We propose a risk-aware evaluation framework for LLM security failures in Banking, Financial Services, and Insurance (BFSI), combining a domain-specific taxonomy of financial harms, an automated multi-round red-teaming pipeline, and an ensemble-based judging protocol. We introduce the Risk-Adjusted Harm Score (RAHS), a risk-sensitive metric that goes beyond success rates by quantifying the operational severity of disclosures, accounting for mitigation signals, and leveraging inter-judge agreement. Across diverse models, we find that higher decoding stochasticity and sustained adaptive interaction not only increase jailbreak success, but also drive systematic escalation toward more severe and operationally actionable financial disclosures. These results expose limitations of single-turn, domain-agnostic security evaluation and motivate risk-sensitive assessment under prolonged adversarial pressure for real-world BFSI deployment.

📄 PDF Abstract BibTeX arXiv:2603.10807

Code (0)

등록된 구현이 없습니다.

Tasks

Red Teaming

Similar Papers 제목 키워드 기반

HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

2024-02-06 · Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 외

Automated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a standardized evaluation framework to rigoro…

Red Teaming

Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique

2024-08-20 · Tej Deep Pala, Vernon Y. H. Toh, Rishabh Bhardwaj, Soujanya Poria

In today's era, where large language models (LLMs) are integrated into numerous real-world applications, ensuring their safety and robustness is crucial for responsible AI usage. Automated red-teaming methods play a key …

AI and SafetyDiversityRed TeamingSafety Alignment

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models

2025-10-09 · Haoran Ou, Kangjie Chen, Xingshuo Han, Gelei Deng 외 arxiv

Large Language Models (LLMs) have been augmented with web search to overcome the limitations of the static knowledge boundary by accessing up-to-date information from the open Internet. While this integration enhances mo…

Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction

2024-09-25 · Jinchuan Zhang, Yan Zhou, Yaxin Liu, Ziming Li 외

Automated red teaming is an effective method for identifying misaligned behaviors in large language models (LLMs). Existing approaches, however, often focus primarily on improving attack success rates while overlooking t…

DiversityRed Teaming

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

2025-06-08 · Ren-Jian Wang, Ke Xue, Zeyu Qin, Ziniu Li 외

Ensuring safety of large language models (LLMs) is important. Red teaming--a systematic approach to identifying adversarial prompts that elicit harmful responses from target LLMs--has emerged as a crucial safety evaluati…

DiversityRed TeamingSentence EmbeddingSentence-Embedding