paper-with-me

홈 › Papers

DynaGuard: A Dynamic Guardian Model With User-Defined Policies

2025-09-02 · Monte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah, Joseph Vincent, Chirag Jain, Melissa Kazemi Rad, C. Bayan Bruss, Ashwinee Panda, Tom Goldstein arxiv

Guardian models play a crucial role in ensuring the safety and ethical behavior of user-facing AI applications by enforcing guardrails and detecting harmful content. While standard guardian models are limited to predefined, static harm categories, we introduce DynaGuard, a suite of dynamic guardian models offering novel flexibility by evaluating text based on user-defined policies, and DynaBench, a dataset for training and evaluating dynamic guardian models. Our models provide both rapid detection of policy violations and a chain-of-thought reasoning option that articulate and justify model outputs. Critically, DynaGuard not only surpasses static models in detection accuracy on traditional safety categories, but is competitive with frontier reasoning models on free-form policy violations, all in a fraction of the time. This makes DynaGuard an critical tool for language model guardrails.

📄 PDF Abstract BibTeX arXiv:2509.02563

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

2026-01-19 · Tassallah Abdullahi, Macton Mgonzo, Mardiyyah Oduwole, Paul Okewunmi 외 arxiv

Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerable to evolving harms, cross-lingual failures, and cultural misalignment.…

Cross-Lingual Transfer

AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior

2026-01-15 · Nadya Abaev, Denis Klimov, Gerard Levinov, David Mimran 외 arxiv

Artificial intelligence (AI) agents are increasingly used in a variety of domains to automate tasks, interact with users, and make decisions based on data inputs. Ensuring that AI agents perform only authorized actions a…

WatchGuardian: Enabling User-Defined Personalized Just-in-Time Intervention on Smartwatch

2025-02-09 · Ying Lei, Yancheng Cao, Will Wang, Yuanzhe Dong 외

While just-in-time interventions (JITIs) have effectively targeted common health behaviors, individuals often have unique needs to intervene in personal undesirable actions that can negatively affect physical, mental, an…

Data AugmentationFew-Shot Learning

Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents

2026-06-01 · Yoshinari Fujinuma, Varun Gangal, Traian Rebedea, Makesh Narsimhan Sreedhar 외 arxiv

Large language model (LLM) agents increasingly rely on reusable skills i.e. documents describing task-specific procedures. However, this introduces a new attack surface for agents to manage. We study two complementary di…

The Rise of Guardians: Fact-checking URL Recommendation to Combat Fake News

2018-07-11 · The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR 2018, Ann Arbor, MI, USA, July 08-12, 2018 2018 7 · Vo, Nguyen; Lee, Kyumin

A large body of research work and efforts have been focused on detecting fake news and building online fact-check systems in order to debunk fake news as soon as possible. Despite the existence of these systems, fake new…

Fact CheckingMisinformationRecommendation Systems