paper-with-me

홈 › Papers

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection

2025-08-10 · Ivan Zhang arxiv

Ensuring LLM alignment is critical to information security as AI models become increasingly widespread and integrated in society. Unfortunately, many defenses against adversarial attacks and jailbreaking on LLMs cannot adapt quickly to new attacks, degrade model responses to benign prompts, or introduce significant barriers to scalable implementation. To mitigate these challenges, we introduce a real-time, self-tuning (RTST) moderator framework to defend against adversarial attacks while maintaining a lightweight training footprint. We empirically evaluate its effectiveness using Google's Gemini models against modern, effective jailbreaks. Our results demonstrate the advantages of an adaptive, minimally intrusive framework for jailbreak defense over traditional fine-tuning or classifier models.

📄 PDF Abstract BibTeX arXiv:2508.07139

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Triaging Content Severity in Online Mental Health Forums

2017-02-22 · Arman Cohan, Sydney Young, Andrew Yates, Nazli Goharian

Mental health forums are online communities where people express their issues and seek help from moderators and other users. In such forums, there are often posts with severe content indicating that the user is in acute …

Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster

2024-06-06 · Agostina Calabrese, Leonardo Neves, Neil Shah, Maarten W. Bos 외

Content moderators play a key role in keeping the conversation on social media healthy. While the high volume of content they need to judge represents a bottleneck to the moderation pipeline, no studies have explored how…

Decision Making

Can Language Model Moderators Improve the Health of Online Discourse?

2023-11-16 · Hyundong Cho, Shuai Liu, Taiwei Shi, Darpan Jain 외

Conversational moderation of online communities is crucial to maintaining civility for a constructive environment, but it is challenging to scale and harmful to moderators. The inclusion of sophisticated natural language…

Language ModelingLanguage ModellingText Generation

WHoW: A Cross-domain Approach for Analysing Conversation Moderation

2024-10-21 · Ming-Bin Chen, Lea Frermann, Jey Han Lau

We propose WHoW, an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who). Using th…

Selective Inference for Time-Varying Effect Moderation

2024-11-24 · Soham Bakshi, Walter Dempsey, Snigdha Panigrahi

Causal effect moderation investigates how the effect of interventions (or treatments) on outcome variables changes based on observed characteristics of individuals, known as potential effect moderators. With advances in …

valid