paper-with-me

홈 › Papers

IndicJR: A Judge-Free Benchmark of Jailbreak Robustness in South Asian Languages

2026-02-18 · Priyaranjan Pattnayak, Sanchari Chowdhuri arxiv

Safety alignment of large language models (LLMs) is mostly evaluated in English and contract-bound, leaving multilingual vulnerabilities understudied. We introduce \textbf{Indic Jailbreak Robustness (IJR)}, a judge-free benchmark for adversarial safety across 12 Indic and South Asian languages (2.1 Billion speakers), covering 45216 prompts in JSON (contract-bound) and Free (naturalistic) tracks. IJR reveals three patterns. (1) Contracts inflate refusals but do not stop jailbreaks: in JSON, LLaMA and Sarvam exceed 0.92 JSR, and in Free all models reach 1.0 with refusals collapsing. (2) English to Indic attacks transfer strongly, with format wrappers often outperforming instruction wrappers. (3) Orthography matters: romanized or mixed inputs reduce JSR under JSON, with correlations to romanization share and tokenization (approx 0.28 to 0.32) indicating systematic effects. Human audits confirm detector reliability, and lite-to-full comparisons preserve conclusions. IJR offers a reproducible multilingual stress test revealing risks hidden by English-only, contract-focused evaluations, especially for South Asian users who frequently code-switch and romanize.

📄 PDF Abstract BibTeX arXiv:2602.16832

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness

2026-02-04 · Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami 외 arxiv

Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate har…

Adversarial Robustness

Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs

2025-08-22 · Yu Yan, Sheng Sun, Zhe Wang, Yijun Lin 외 arxiv

With the development of Large Language Models (LLMs), numerous efforts have revealed their vulnerabilities to jailbreak attacks. Although these studies have driven the progress in LLMs' safety alignment, it remains uncle…

D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting

2026-05-31 · Huanli Gong, Zhipeng Wei, Yu Fu, Haz Sameen Shahgir 외 arxiv

Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge models to iteratively refine prompts toward harmful goals. Existing defenses larg…

TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors

2026-09-04 · Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho arxiv

Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful prompts with varying threat implicitness. We introduce TIER, a Threat Implicitness Benchmark for behavioral safety e…

JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

2024-10-11 · Fan Liu, Yue Feng, Zhao Xu, Lixin Su 외

Despite advancements in enhancing LLM safety against jailbreak attacks, evaluating LLM defenses remains a challenge, with current methods often lacking explainability and generalization to complex scenarios, leading to i…