paper-with-me

홈 › Papers

CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation

2026-02-09 · Chaeyun Kim, YongTaek Lim, Kihyun Kim, Junghwan Kim, Minwoo Kim arxiv

Existing red-teaming benchmarks, when adapted to new languages via direct translation, fail to capture socio-technical vulnerabilities rooted in local culture and law, creating a critical blind spot in LLM safety evaluation. To address this gap, we introduce CAGE (Culturally Adaptive Generation), a framework that systematically adapts the adversarial intent of proven red-teaming prompts to new cultural contexts. At the core of CAGE is the Semantic Mold, a novel approach that disentangles a prompt's adversarial structure from its cultural content. This approach enables the modeling of realistic, localized threats rather than testing for simple jailbreaks. As a representative example, we demonstrate our framework by creating KoRSET, a Korean benchmark, which proves more effective at revealing vulnerabilities than direct translation baselines. CAGE offers a scalable solution for developing meaningful, context-aware safety benchmarks across diverse cultures. Our dataset and evaluation rubrics are publicly available at https://github.com/selectstar-ai/CAGE-paper. (WARNING: This paper contains model outputs that can be offensive in nature.)

📄 PDF Abstract BibTeX arXiv:2602.20170

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

2026-03-13 · Chenlong Yin, Runpeng Geng, Yanting Wang, Jinyuan Jia arxiv

Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents. Although many defenses have been proposed, their robustness against adaptive attacks remains insufficiently ev…

Reinforcement LearningRed Teaming

SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset

2026-06-29 · Ariel Gjaci, Antonio Sgorbissa, Vittorio Murino arxiv

Recent co-speech gesture generation methods often overlook cultural differences, limiting their effectiveness in human-agent interaction. Moreover, culture-conditioned models are rarely evaluated under speaker-disjoint s…

Domain GeneralizationGesture GenerationMotion Synthesis

IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection

2026-05-12 · Chia-Pei, Chen, Kentaroh Toyoda, Anita Lai 외 arxiv

Web-browsing AI agents are increasingly deployed in enterprise settings under strict whitelists of approved domains, yet adversaries can still influence them by embedding hidden instructions in the HTML pages those domai…

Signals Are Not States: Neuro-Symbolic Safeguards for Culturally Aware Classroom AI

2026-03-24 · Sina Bagheri Nezhad arxiv

Classroom AI systems increasingly infer high-level educational states such as engagement, confusion, collaboration, participation, and instructional quality from multimodal and linguistic signals. In multicultural and mu…

AT-Drone: Benchmarking Adaptive Teaming in Multi-Drone Pursuit

2025-02-13 · Yang Li, Junfan Chen, Feng Xue, Jiabin Qiu 외

Adaptive teaming-the capability of agents to effectively collaborate with unfamiliar teammates without prior coordination-is widely explored in virtual video games but overlooked in real-world multi-robot contexts. Yet, …

BenchmarkingEdge-computing