paper-with-me

Papers

Swiss Cheese Model for AI Safety: A Taxonomy and Reference Architecture for Multi-Layered Guardrails of Foundation Model Based Agents

2024-08-05 · Md Shamsujjoha, Qinghua Lu, Dehai Zhao, Liming Zhu

Foundation Model (FM)-based agents are revolutionizing application development across various domains. However, their rapidly growing capabilities and autonomy have raised significant concerns about AI safety. Researchers are exploring better ways to design guardrails to ensure that the runtime behavior of FM-based agents remains within specific boundaries. Nevertheless, designing effective runtime guardrails is challenging due to the agents' autonomous and non-deterministic behavior. The involvement of multiple pipeline stages and agent artifacts, such as goals, plans, tools, at runtime further complicates these issues. Addressing these challenges at runtime requires multi-layered guardrails that operate effectively at various levels of the agent architecture. Therefore, in this paper, based on the results of a systematic literature review, we present a comprehensive taxonomy of runtime guardrails for FM-based agents to identify the key quality attributes for guardrails and design dimensions. Inspired by the Swiss Cheese Model, we also propose a reference architecture for designing multi-layered runtime guardrails for FM-based agents, which includes three dimensions: quality attributes, pipelines, and artifacts. The proposed taxonomy and reference architecture provide concrete and robust guidance for researchers and practitioners to build AI-safety-by-design from a software architecture perspective.

📄 PDF Abstract BibTeX arXiv:2408.02205

Code (1)

dishacse/Publication-Resources/tree/main/2025%20ICSA 공식 구현

Tasks

modelSystematic Literature Review

Similar Papers 제목 키워드 기반

SwissCheese at SemEval-2016 Task 4: Sentiment Classification Using an Ensemble of Convolutional Neural Networks with Distant Supervision

2016-06-01 · SEMEVAL 2016 6 · Jan Deriu, Maurice Gonzenbach, Fatih Uzdilli, Aurelien Lucchi 외
General ClassificationSentence EmbeddingSentiment AnalysisSentiment Classification+1

Model Spec Midtraining: Improving How Alignment Training Generalizes

2026-05-03 · Chloe Li, Nevan Wichers, Sara Price, Samuel Marks 외 arxiv

Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning -- training on demonstrations of spec-aligned b…

Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits

2026-01-28 · Zelong Sun, Jiahui Wu, Ying Ba, Dong Jing 외 arxiv

As social media platforms proliferate, users increasingly demand intuitive ways to create diverse, high-quality portrait collections. In this work, we introduce Portrait Collection Generation (PCG), a novel task that gen…

What are the odds? Risk and uncertainty about AI existential risk

2025-10-27 · Marco Grossi arxiv

This work is a commentary of the article \href{https://doi.org/10.18716/ojs/phai/2025.2801}{AI Survival Stories: a Taxonomic Analysis of AI Existential Risk} by Cappelen, Goldstein, and Hawthorne. It is not just a commen…

Vitamin K content of cheese, yoghurt and meat products in Australia

2022-03-22 · Eleanor Dunlop, Jette Jakobsen, Marie Bagge Jensen, Jayashree Arcot 외

Vitamin K is vital for normal blood coagulation, and may influence bone, neurological and vascular health. Data on the vitamin K content of Australian foods are limited, preventing estimation of vitamin K intakes in the …