paper-with-me

홈 › Papers

A Grounded Observer Framework for Establishing Guardrails for Foundation Models in Socially Sensitive Domains

2024-12-23 · Rebecca Ramnauth, Dražen Brščić, Brian Scassellati

As foundation models increasingly permeate sensitive domains such as healthcare, finance, and mental health, ensuring their behavior meets desired outcomes and social expectations becomes critical. Given the complexities of these high-dimensional models, traditional techniques for constraining agent behavior, which typically rely on low-dimensional, discrete state and action spaces, cannot be directly applied. Drawing inspiration from robotic action selection techniques, we propose the grounded observer framework for constraining foundation model behavior that offers both behavioral guarantees and real-time variability. This method leverages real-time assessment of low-level behavioral characteristics to dynamically adjust model actions and provide contextual feedback. To demonstrate this, we develop a system capable of sustaining contextually appropriate, casual conversations ("small talk"), which we then apply to a robot for novel, unscripted interactions with humans. Finally, we discuss potential applications of the framework for other social contexts and areas for further research.

📄 PDF Abstract BibTeX arXiv:2412.18639

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains

2026-05-19 · Rebecca Ramnauth, Drazen Brscic, Brian Scassellati arxiv

Foundation models are increasingly deployed in socially sensitive domains such as education, mental health, and caregiving, where failures are often cumulative and context-dependent. Existing guardrail approaches -- rang…

A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace

2026-07-22 · Kathrin Paimann, Elizangela Valarini, Sebastian Juhl arxiv

As AI agents become integral to business workflows, establishing guiding user experience (UX) principles is crucial for ensuring user trust and successful adoption. To address this, our study uses a multi-method approach…

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

2026-05-29 · Yan Wang, Zhixuan Chu, Zihao Xue, Zhen Bi 외 arxiv

Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do not always lead to faithful enforcement: a model may recognize a har…

Alpay Algebra III: Observer-Coupled Collapse and the Temporal Drift of Identity

2025-05-26 · Faruk Alpay

This paper introduces a formal framework for modeling observer-dependent collapse dynamics and temporal identity drift within artificial and mathematical systems, grounded entirely in the symbolic foundations of Alpay Al…

Formal Logic

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

2026-04-17 · Hua-Rong Chu, Kuan-Chun Wang, Yao-Te Huang arxiv

Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). However, existing research lacks consideration of nuances across linguist…