paper-with-me

Papers

Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems

2025-10-15 · Karthik Avinash, Nikhil Pareek, Rishav Hada arxiv

The increasing deployment of Large Language Models (LLMs) across enterprise and mission-critical domains has underscored the urgent need for robust guardrailing systems that ensure safety, reliability, and compliance. Existing solutions often struggle with real-time oversight, multi-modal data handling, and explainability -- limitations that hinder their adoption in regulated environments. Existing guardrails largely operate in isolation, focused on text alone making them inadequate for multi-modal, production-scale environments. We introduce Protect, natively multi-modal guardrailing model designed to operate seamlessly across text, image, and audio inputs, designed for enterprise-grade deployment. Protect integrates fine-tuned, category-specific adapters trained via Low-Rank Adaptation (LoRA) on an extensive, multi-modal dataset covering four safety dimensions: toxicity, sexism, data privacy, and prompt injection. Our teacher-assisted annotation pipeline leverages reasoning and explanation traces to generate high-fidelity, context-aware labels across modalities. Experimental results demonstrate state-of-the-art performance across all safety dimensions, surpassing existing open and proprietary models such as WildGuard, LlamaGuard-4, and GPT-4.1. Protect establishes a strong foundation for trustworthy, auditable, and production-ready safety systems capable of operating across text, image, and audio modalities.

📄 PDF Abstract BibTeX arXiv:2510.13351

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Authenticated Workflows: A Systems Approach to Protecting Agentic AI

2026-02-11 · Mohan Rajagopalan, Vinay Rao arxiv

Agentic AI systems automate enterprise workflows but existing defenses--guardrails, semantic filters--are probabilistic and routinely bypassed. We introduce authenticated workflows, the first complete trust layer for ent…

Trustworthy AI Inference Systems: An Industry Research View

2020-08-10 · Rosario Cammarota, Matthias Schunter, Anand Rajan, Fabian Boemer 외

In this work, we provide an industry research view for approaching the design, deployment, and operation of trustworthy Artificial Intelligence (AI) inference systems. Such systems provide customers with timely, informed…

AI Governance Control Stack for Operational Stability: Achieving Hardened Governance in AI Systems

2026-03-12 · Horatio Morgan arxiv

Artificial intelligence systems are increasingly embedded in high-stakes decision environments, yet many governance approaches focus primarily on policy guidance rather than operational stability mechanisms. As AI deploy…

SynRAG: A Large Language Model Framework for Executable Query Generation in Heterogeneous SIEM System

2025-12-31 · Md Hasan Saju, Austin Page, Akramul Azim, Jeff Gardiner 외 arxiv

Security Information and Event Management (SIEM) systems are essential for large enterprises to monitor their IT infrastructure by ingesting and analyzing millions of logs and events daily. Security Operations Center (SO…

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

2026-01-17 · YenTing Lee, Keerthi Koneru, Zahra Moslemi, Sheethal Kumar 외 arxiv

Evaluating large language model (LLM)-based multi-agent systems remains a critical challenge, as these systems must exhibit reliable coordination, transparent decision-making, and verifiable performance across evolving t…