paper-with-me

Papers

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety

2026-04-20 · Marcello Galisai, Susanna Cifani, Francesco Giarrusso, Piercosma Bisconti, Matteo Prandi, Federico Pierucci, Federico Sartore, Daniele Nardi arxiv

The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewrites the same objectives through humanities-style transformations while preserving intent. This extends literature on Adversarial Poetry and Adversarial Tales from single jailbreak operators to a broader benchmark family of stylistic obfuscation and goal concealment. In the benchmark results reported here, the original attacks record 3.84% attack success rate (ASR), while transformed methods range from 36.8% to 65.0%, yielding 55.75% overall ASR across 31 frontier models. Under a European Union AI Act Code-of-Practice-inspired systemic-risk lens, Chemical, biological, radiological and nuclear (CBRN) is the highest bucket. Taken together, this lack of stylistic robustness suggests that current safety techniques suffer from weak generalization: deep understanding of 'non-maleficence' remains a central unresolved problem in frontier model safety.

📄 PDF Abstract BibTeX arXiv:2604.18487

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarking Large Language Models on Reference Extraction and Parsing in the Social Sciences and Humanities

2026-03-13 · Yurui Zhu, Giovanni Colavizza, Matteo Romanello arxiv

Bibliographic reference extraction and parsing are foundational for citation indexing, linking, and downstream scholarly knowledge-graph construction. However, most established evaluations focus on clean, English, end-of…

Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation

2026-05-07 · Siyuan Li, Aodu Wulianghai, Xi Lin, Xibin Yuan 외 arxiv

The increasing prevalence of Large Language Models (LLMs) in content creation has made distinguishing human-written textual content from LLM-generated counterparts a critical task for multimedia moderation. Existing dete…

Modeling intra-textual variation with entropy and surprisal: topical vs. stylistic patterns

2017-08-01 · WS 2017 8 · Stefania Degaetano-Ortlieb, Elke Teich

We present a data-driven approach to investigate intra-textual variation by combining entropy and surprisal. With this approach we detect linguistic variation based on phrasal lexico-grammatical patterns across sections …

Articles

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

2026-08-03 · Fengxian Ji, Yuke Li, Jingpu Yang, Juanfan Wu 외 arxiv

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unifi…

Opportunities for Persian Digital Humanities Research with Artificial Intelligence Language Models; Case Study: Forough Farrokhzad

2024-05-10 · Arash Rasti Meymandi, Zahra Hosseini, Sina Davari, Abolfazl Moshiri 외

This study explores the integration of advanced Natural Language Processing (NLP) and Artificial Intelligence (AI) techniques to analyze and interpret Persian literature, focusing on the poetry of Forough Farrokhzad. Uti…