paper-with-me

홈 › Papers

Retrieval-Augmented Guardrails for AI-Drafted Patient-Portal Messages: Error Taxonomy Construction and Large-Scale Evaluation

2025-09-26 · Wenyuan Chen, Fateme Nateghi Haredasht, Kameron C. Black, Francois Grolleau, Emily Alsentzer, Jonathan H. Chen, Stephen P. Ma arxiv

Asynchronous patient-clinician messaging via EHR portals is a growing source of clinician workload, prompting interest in large language models (LLMs) to assist with draft responses. However, LLM outputs may contain clinical inaccuracies, omissions, or tone mismatches, making robust evaluation essential. Our contributions are threefold: (1) we introduce a clinically grounded error ontology comprising 5 domains and 59 granular error codes, developed through inductive coding and expert adjudication; (2) we develop a retrieval-augmented evaluation pipeline (RAEC) that leverages semantically similar historical message-response pairs to improve judgment quality; and (3) we provide a two-stage prompting architecture using DSPy to enable scalable, interpretable, and hierarchical error detection. Our approach assesses the quality of drafts both in isolation and with reference to similar past message-response pairs retrieved from institutional archives. Using a two-stage DSPy pipeline, we compared baseline and reference-enhanced evaluations on over 1,500 patient messages. Retrieval context improved error identification in domains such as clinical completeness and workflow appropriateness. Human validation on 100 messages demonstrated superior agreement (concordance = 50% vs. 33%) and performance (F1 = 0.500 vs. 0.256) of context-enhanced labels vs. baseline, supporting the use of our RAEC pipeline as AI guardrails for patient messaging.

📄 PDF Abstract BibTeX arXiv:2509.22565

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting

2026-01-16 · Parker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe 외 arxiv

Large language models (LLMs) show promise in drafting responses to patient portal messages, yet their integration into clinical workflows raises various concerns, including whether they would actually save clinicians tim…

Lab-AI: Using Retrieval Augmentation to Enhance Language Models for Personalized Lab Test Interpretation in Clinical Medicine

2024-09-16 · Xiaoyu Wang, Haoyong Ouyang, Balu Bhasuran, Xiao Luo 외

Accurate interpretation of lab results is crucial in clinical medicine, yet most patient portals use universal normal ranges, ignoring conditional factors like age and gender. This study introduces Lab-AI, an interactive…

Language ModelingLanguage ModellingRAGRetrieval+1

RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts

2025-10-06 · Yining She, Daniel W. Peterson, Marianne Menglin Liu, Vikas Upadhyay 외 arxiv

With the increasing adoption of large language models (LLMs), ensuring the safety of LLM systems has become a pressing concern. External LLM-based guardrail models have emerged as a popular solution to screen unsafe inpu…

In-Context Learning for Preserving Patient Privacy: A Framework for Synthesizing Realistic Patient Portal Messages

2024-11-10 · Joseph Gatto, Parker Seegmiller, Timothy E. Burdick, Sarah Masud Preum

Since the COVID-19 pandemic, clinicians have seen a large and sustained influx in patient portal messages, significantly contributing to clinician burnout. To the best of our knowledge, there are no large-scale public pa…

De-identificationIn-Context LearningPrivacy PreservingText Generation

Retrieval-Augmented Generation for Natural Language Art Provenance Searches in the Getty Provenance Index

2025-08-26 · Mathew Henrickson arxiv

This research presents a Retrieval-Augmented Generation (RAG) framework for art provenance studies, focusing on the Getty Provenance Index. Provenance research establishes the ownership history of artworks, which is esse…

Semantic Retrieval