paper-with-me

홈 › Papers

Safeguarding AI Agents: Developing and Analyzing Safety Architectures

2024-09-03 · Ishaan Domkundwar, Mukunda N S, Ishaan Bhola, Riddhik Kochhar

AI agents, specifically powered by large language models, have demonstrated exceptional capabilities in various applications where precision and efficacy are necessary. However, these agents come with inherent risks, including the potential for unsafe or biased actions, vulnerability to adversarial attacks, lack of transparency, and tendency to generate hallucinations. As AI agents become more prevalent in critical sectors of the industry, the implementation of effective safety protocols becomes increasingly important. This paper addresses the critical need for safety measures in AI systems, especially ones that collaborate with human teams. We propose and evaluate three frameworks to enhance safety protocols in AI agent systems: an LLM-powered input-output filter, a safety agent integrated within the system, and a hierarchical delegation-based system with embedded safety checks. Our methodology involves implementing these frameworks and testing them against a set of unsafe agentic use cases, providing a comprehensive evaluation of their effectiveness in mitigating risks associated with AI agent deployment. We conclude that these frameworks can significantly strengthen the safety and security of AI agent systems, minimizing potential harmful actions or outputs. Our work contributes to the ongoing effort to create safe and reliable AI applications, particularly in automated operations, and provides a foundation for developing robust guardrails to ensure the responsible use of AI agents in real-world applications.

📄 PDF Abstract BibTeX arXiv:2409.03793

Code (0)

등록된 구현이 없습니다.

Tasks

AI Agent

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Prioritizing Safeguarding Over Autonomy: Risks of LLM Agents for Science

2024-02-06 · Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan 외

Intelligent agents powered by large language models (LLMs) have demonstrated substantial promise in autonomously conducting experiments and facilitating scientific discoveries across various disciplines. While their capa…

TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems

2026-03-16 · Kai Wang, Biaojie Zeng, Zeming Wei, Chang Jin 외 arxiv

With the rapid development of LLM-based multi-agent systems (MAS), their significant safety and security concerns have emerged, which introduce novel risks going beyond single agents or LLMs. Despite attempts to address …

Safe Sliding Mode Controllers for Nonlinear Uncertain Systems

2024-06-06 · Yazdan Batmani, Mohammadreza Davoodi

In this study, we present a novel sliding mode safety-critical controller designed to address both stability and safety concerns in a class of nonlinear uncertain systems. The controller features two feedback loops: an i…

Provably Safe Reinforcement Learning from Analytic Gradients

2025-06-02 · Tim Walter, Hannah Markgraf, Jonathan Külz, Matthias Althoff

Deploying autonomous robots in safety-critical applications requires safety guarantees. Provably safe reinforcement learning is an active field of research which aims to provide such guarantees using safeguards. These sa…

reinforcement-learningReinforcement LearningSafe Reinforcement Learning

Towards Safe Autonomous Driving: A Real-Time Motion Planning Algorithm on Embedded Hardware

2026-01-07 · Korbinian Moller, Glenn Johannes Tungka, Lucas Jürgens, Johannes Betz arxiv

Ensuring the functional safety of Autonomous Vehicles (AVs) requires motion planning modules that not only operate within strict real-time constraints but also maintain controllability in case of system faults. Existing …

Trajectory PlanningAutonomous VehiclesAutonomous DrivingMotion Planning