paper-with-me

홈 › Papers

Over-Relying on Reliance: Towards Realistic Evaluations of AI-Based Clinical Decision Support

2025-04-10 · Venkatesh Sivaraman, Katelyn Morrison, Will Epperson, Adam Perer

As AI-based clinical decision support (AI-CDS) is introduced in more and more aspects of healthcare services, HCI research plays an increasingly important role in designing for complementarity between AI and clinicians. However, current evaluations of AI-CDS often fail to capture when AI is and is not useful to clinicians. This position paper reflects on our work and influential AI-CDS literature to advocate for moving beyond evaluation metrics like Trust, Reliance, Acceptance, and Performance on the AI's task (what we term the "trap" of human-AI collaboration). Although these metrics can be meaningful in some simple scenarios, we argue that optimizing for them ignores important ways that AI falls short of clinical benefit, as well as ways that clinicians successfully use AI. As the fields of HCI and AI in healthcare develop new ways to design and evaluate CDS tools, we call on the community to prioritize ecologically valid, domain-appropriate study setups that measure the emergent forms of value that AI can bring to healthcare professionals.

📄 PDF Abstract BibTeX arXiv:2504.07423

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Similar Papers 제목 키워드 기반

Can SAEs reveal and mitigate racial biases of LLMs in healthcare?

2025-10-31 · Hiba Ahsan, Byron C. Wallace arxiv

LLMs are increasingly being used in healthcare. This promises to free physicians from drudgery, enabling better care to be delivered at scale. But the use of LLMs in this space also brings risks; for example, such models…

Applying and Evaluating Large Language Models in Mental Health Care: A Scoping Review of Human-Assessed Generative Tasks

2024-08-21 · Yining Hua, Hongbin Na, Zehan Li, Fenglin Liu 외

Large language models (LLMs) are emerging as promising tools for mental health care, offering scalable support through their ability to generate human-like responses. However, the effectiveness of these models in clinica…

ArticlesFairness

Testing Generalizability in Causal Inference

2024-11-05 · Daniel de Vassimon Manela, Linying Yang, Robin J. Evans

Ensuring robust model performance in diverse real-world scenarios requires addressing generalizability across domains with covariate shifts. However, no formal procedure exists for statistically evaluating generalizabili…

Causal InferenceDecision Making

AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

2026-05-27 · Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig 외 arxiv

AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration requires understanding when, why, and how humans decide to rely on AI.…

Question Answering

AgentEHR: Advancing Autonomous Clinical Decision-Making via Retrospective Summarization

2026-01-20 · Yusheng Liao, Chuan Xuan, Yutong Cai, Lina Yang 외 arxiv

Large Language Models have demonstrated profound utility in the medical domain. However, their application to autonomous Electronic Health Records~(EHRs) navigation remains constrained by a reliance on curated inputs and…