paper-with-me

홈 › Papers

RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions

2026-01-19 · Drishti Goel, Jeongah Lee, Qiuyue Joy Zhong, Violeta J. Rodriguez, Daniel S. Brown, Ravi Karkar, Dong Whi Yoo, Koustuv Saha arxiv

Caregivers seeking AI-mediated support express complex needs -- information-seeking, emotional validation, and distress cues -- that warrant careful evaluation of response safety and appropriateness. Existing AI evaluation frameworks, primarily focused on general risks (toxicity, hallucinations, policy violations, etc), may not adequately capture the nuanced risks of LLM-responses in caregiving-contexts. We introduce RubRIX (Rubric-based Risk Index), a theory-driven, clinician-validated framework for evaluating risks in LLM caregiving responses. Grounded in the Elements of an Ethic of Care, RubRIX operationalizes five empirically-derived risk dimensions: Inattention, Bias & Stigma, Information Inaccuracy, Uncritical Affirmation, and Epistemic Arrogance. We evaluate six state-of-the-art LLMs on over 20,000 caregiver queries from Reddit and ALZConnected. Rubric-guided refinement consistently reduced risk-components by 45-98% after one iteration across models. This work contributes a methodological approach for developing domain-sensitive, user-centered evaluation frameworks for high-burden contexts. Our findings highlight the importance of domain-sensitive, interactional risk evaluation for the responsible deployment of LLMs in caregiving support contexts. We release benchmark datasets to enable future research on contextual risk evaluation in AI-mediated support.

📄 PDF Abstract BibTeX arXiv:2601.13235

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mapping Caregiver Needs to AI Chatbot Design: Strengths and Gaps in Mental Health Support for Alzheimer's and Dementia Caregivers

2025-06-18 · Jiayue Melissa Shi, Dong Whi Yoo, Keran Wang, Violeta J. Rodriguez 외

Family caregivers of individuals with Alzheimer's Disease and Related Dementia (AD/ADRD) face significant emotional and logistical challenges that place them at heightened risk for stress, anxiety, and depression. Althou…

Chatbot

Autorubric: Unifying Rubric-based LLM Evaluation

2026-02-13 · Delip Rao, Chris Callison-Burch arxiv

Techniques for reliable rubric-based LLM evaluation -- ensemble judging, bias mitigation, few-shot calibration -- are scattered across papers with inconsistent terminology and partial implementations. We introduce Autoru…

Why Are We Lonely? Leveraging LLMs to Measure and Understand Loneliness in Caregivers and Non-caregivers

2026-04-09 · Michelle Damin Kim, Ellie S. Paek, Yufen Lin, Emily Mroz 외 arxiv

This paper presents an LLM-driven approach for constructing diverse social media datasets to measure and compare loneliness in the caregiver and non-caregiver populations. We introduce an expert-developed loneliness eval…

Risks of AI-driven product development and strategies for their mitigation

2025-05-28 · Jan Göpfert, Jann M. Weinand, Patrick Kuckertz, Noah Pflugradt 외

Humanity is progressing towards automated product development, a trend that promises faster creation of better products and thus the acceleration of technological progress. However, increasing reliance on non-human agent…

Online Rubrics Elicitation from Pairwise Comparisons

2025-10-08 · MohammadHossein Rezaei, Robert Vacareanu, Zihao Wang, Clinton Wang 외 arxiv

Rubrics provide a flexible way to train LLMs on open-ended long-form answers where verifiable rewards are not applicable and human preferences provide coarse signals. Prior work shows that reinforcement learning with rub…

Reinforcement Learning