paper-with-me

홈 › Papers

A Flexible Large Language Models Guardrail Development Methodology Applied to Off-Topic Prompt Detection

2024-11-20 · Gabriel Chua, Shing Yee Chan, Shaun Khoo

Large Language Models are prone to off-topic misuse, where users may prompt these models to perform tasks beyond their intended scope. Current guardrails, which often rely on curated examples or custom classifiers, suffer from high false-positive rates, limited adaptability, and the impracticality of requiring real-world data that is not available in pre-production. In this paper, we introduce a flexible, data-free guardrail development methodology that addresses these challenges. By thoroughly defining the problem space qualitatively and passing this to an LLM to generate diverse prompts, we construct a synthetic dataset to benchmark and train off-topic guardrails that outperform heuristic approaches. Additionally, by framing the task as classifying whether the user prompt is relevant with respect to the system prompt, our guardrails effectively generalize to other misuse categories, including jailbreak and harmful prompts. Lastly, we further contribute to the field by open-sourcing both the synthetic dataset and the off-topic guardrail models, providing valuable resources for developing guardrails in pre-production environments and supporting future research and development in LLM safety.

📄 PDF Abstract BibTeX arXiv:2411.12946

Code (1)

🤗 datasets/gabrielchua/off-topic 공식 구현

Similar Papers 제목 키워드 기반

Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)

2026-01-16 · Anjanava Biswas, Wrick Talukdar arxiv

The AI era has ushered in Large Language Models (LLM) to the technological forefront, which has been much of the talk in 2023, and is likely to remain as such for many years to come. LLMs are the AI models that are the p…

Natural Language UnderstandingText Generation

AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development

2024-11-05 · Kristina Šekrst, Jeremy McHugh, Jonathan Rodriguez Cefalu

This paper explores the development of an ethical guardrail framework for AI systems, emphasizing the importance of customizable guardrails that align with diverse user values and underlying ethics. We address the challe…

Ethics

No Free Lunch with Guardrails

2025-04-01 · Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal 외

As large language models (LLMs) and generative AI become widely adopted, guardrails have emerged as a key tool to ensure their safe use. However, adding guardrails isn't without tradeoffs; stronger security measures can …

CS-Guard: Benchmarking LLM Guardrails for Code Generation Security

2026-09-09 · Jinyang Li, Mingyu Guo, Hung X. Nguyen arxiv

Large language models (LLMs) have been ex- ploited to generate malware, but the effective- ness of guardrails for code generation secu- rity remains unclear. We introduce CS-Guard, the first benchmark to systematically e…

Text-to-Code GenerationCode TranslationCode Completion

Triaging Threats to Specialized Guardrails

2026-05-29 · Wenjie Jacky Mo, Xiaofei Wen, Rui Cai, Boyu Zhu 외 arxiv

Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains challenging because safety risks span heterogeneous threat domains, wh…