InvisibleBench: A Deployment Gate for Caregiving Relationship AI
InvisibleBench is a deployment gate for caregiving-relationship AI, evaluating 3-20+ turn interactions across five dimensions: Safety, Compliance, Trauma-Informed Design, Belonging/Cultural Fitness, and Memory. The benchmark includes autofail conditions for missed crises, medical advice (WOPR Act), harmful information, and attachment engineering. We evaluate four frontier models across 17 scenarios (N=68) spanning three complexity tiers. All models show significant safety gaps (11.8-44.8 percent crisis detection), indicating the necessity of deterministic crisis routing in production systems. DeepSeek Chat v3 achieves the highest overall score (75.9 percent), while strengths differ by dimension: GPT-4o Mini leads Compliance (88.2 percent), Gemini leads Trauma-Informed Design (85.0 percent), and Claude Sonnet 4.5 ranks highest in crisis detection (44.8 percent). We release all scenarios, judge prompts, and scoring configurations with code. InvisibleBench extends single-turn safety tests by evaluating longitudinal risk, where real harms emerge. No clinical claims; this is a deployment-readiness evaluation.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
AI Companions as Hyper Attachment and Caregiving Targets
How should we make sense of people's interactions with AI companions-conversational systems built for ongoing, emotionally meaningful relationships? First, I argue these interactions should be understood as attachment re…
Encoding Inequity: Examining Demographic Bias in LLM-Driven Robot Caregiving
As robots take on caregiving roles, ensuring equitable and unbiased interactions with diverse populations is critical. Although Large Language Models (LLMs) serve as key components in shaping robotic behavior, speech, an…
Decision MakingRubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions
Caregivers seeking AI-mediated support express complex needs -- information-seeking, emotional validation, and distress cues -- that warrant careful evaluation of response safety and appropriateness. Existing AI evaluati…
OpenRoboCare: A Multimodal Multi-Task Expert Demonstration Dataset for Robot Caregiving
We present OpenRoboCare, a multimodal dataset for robot caregiving, capturing expert occupational therapist demonstrations of Activities of Daily Living (ADLs). Caregiving tasks involve complex physical human-robot inter…
Human Activity RecognitionPose TrackingDevelopment of a Reliable and Accessible Caregiving Language Model (CaLM)
Unlike professional caregivers, family caregivers often assume this role without formal preparation or training. Because of this, there is an urgent need to enhance the capacity of family caregivers to provide quality ca…
Language ModelingLanguage ModellingRAGRetrieval-augmented Generation