paper-with-me

홈 › Papers

FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation

2026-02-27 · Zhihao Ding, Jinming Li, Ze Lu, Jieming Shi arxiv

Ensuring the safety of LLM-generated content is essential for real-world deployment. Most existing guardrail models formulate moderation as a fixed binary classification task, implicitly assuming a fixed definition of harmfulness. In practice, enforcement strictness - how conservatively harmfulness is defined and enforced - varies across platforms and evolves over time, making binary moderators brittle under shifting requirements. We first introduce FlexBench, a strictness-adaptive LLM moderation benchmark that enables controlled evaluation under multiple strictness regimes. Experiments on FlexBench reveal substantial cross-strictness inconsistency in existing moderators: models that perform well under one regime can degrade substantially under others, limiting their practical usability. To address this, we propose FlexGuard, an LLM-based moderator that outputs a calibrated continuous risk score reflecting risk severity and supports strictness-specific decisions via thresholding. We train FlexGuard via risk-alignment optimization to improve score-severity consistency and provide practical threshold selection strategies to adapt to target strictness at deployment. Experiments on FlexBench and public benchmarks demonstrate that FlexGuard achieves higher moderation accuracy and substantially improved robustness under varying strictness. We release the source code and data to support reproducibility.

📄 PDF Abstract BibTeX arXiv:2602.23636

Code (0)

등록된 구현이 없습니다.

Tasks

Binary Classification

Similar Papers 제목 키워드 기반

FLARE: Adaptive Multi-Dimensional Reputation for Robust Client Reliability in Federated Learning

2025-11-18 · Abolfazl Younesi, Leon Kiss, Zahra Najafabadi Samani, Juan Aznar Poveda 외 arxiv

Federated learning (FL) enables collaborative model training while preserving data privacy. However, it remains vulnerable to malicious clients who compromise model integrity through Byzantine attacks, data poisoning, or…

Binary ClassificationFederated Learning

TRUST-SCF: Transformer-based Risk Understanding and Scoring for Transactional Supply Chain Finance

2026-06-06 · Mohammadamin Davoodabadi, Amirabbas Shakeri arxiv

Supply Chain Finance (SCF) and LendTech platforms need credit scoring systems that respond to evolving transaction behavior, repayment delays, and active exposure. We propose TRUST-SCF, a transformer-based framework for …

AI-Driven IRM: Transforming insider risk management with adaptive scoring and LLM-based threat detection

2025-05-01 · Lokesh Koli, Shubham Kalra, Rohan Thakur, Anas Saifi 외

Insider threats pose a significant challenge to organizational security, often evading traditional rule-based detection systems due to their subtlety and contextual nature. This paper presents an AI-powered Insider Risk …

Anomaly DetectionFederated LearningManagement

Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean

2026-04-04 · Aleksandr Meshkov arxiv

Existing evaluation methods for LLM-based AI systems, such as LLM-as-a-Judge, verdict systems, and NLI, do not always align well with human assessment because they cannot adapt their strictness to the application domain.…

CVPL: A Geometric Framework for Post-Hoc Linkage Risk Assessment in Protected Tabular Data

2026-02-11 · Valery Khvatov, Alexey Neyman arxiv

Formal privacy metrics provide compliance-oriented guarantees but often fail to quantify actual linkability in released datasets. We introduce CVPL (Cluster-Vector-Projection Linkage), a geometric framework for post-hoc …