paper-with-me

홈 › Papers

When AI Gets it Wrong: Reliability and Risk in AI-Assisted Medication Decision Systems

2026-04-01 · Khalid Adnan Alsayed arxiv

Artificial intelligence (AI) systems are increasingly integrated into healthcare and pharmacy workflows, supporting tasks such as medication recommendations, dosage determination, and drug interaction detection. While these systems often demonstrate strong performance under standard evaluation metrics, their reliability in real-world decision-making remains insufficiently understood. In high-risk domains such as medication management, even a single incorrect recommendation can result in severe patient harm. This paper examines the reliability of AI-assisted medication systems by focusing on system failures and their potential clinical consequences. Rather than evaluating performance solely through aggregate metrics, this work shifts attention towards how errors occur and what happens when AI systems produce incorrect outputs. Through a series of controlled, simulated scenarios involving drug interactions and dosage decisions, we analyse different types of system failures, including missed interactions, incorrect risk flagging, and inappropriate dosage recommendations. The findings highlight that AI errors in medication-related contexts can lead to adverse drug reactions, ineffective treatment, or delayed care, particularly when systems are used without sufficient human oversight. Furthermore, the paper discusses the risks of over-reliance on AI recommendations and the challenges posed by limited transparency in decision-making processes. This work contributes a reliability-focused perspective on AI evaluation in healthcare, emphasising the importance of understanding failure behavior and real-world impact. It highlights the need to complement traditional performance metrics with risk-aware evaluation approaches, particularly in safety-critical domains such as pharmacy practice.

📄 PDF Abstract BibTeX arXiv:2604.01449

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Budgeted Act-or-Defer Multi-Agent LLM Deliberation with Local Reliability Bounds

2026-06-28 · Mengdie Flora Wang, Haochen Xie, Guanghui Wang, Devin Zhang 외 arxiv

Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review. We formulate this as budge…

Decision Making

What is the Sharpe Ratio, and how can everyone get it wrong?

2018-02-13

The Sharpe ratio is the most widely used risk metric in the quantitative finance community - amazingly, essentially everyone gets it wrong. In this note, we will make a quixotic effort to rectify the situation.

ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

2026-06-13 · Rahul Suresh Babu, Laxmipriya Ganesh Iyer arxiv

Tool-augmented large language model agents increasingly operate over large tool libraries, but existing evaluations often focus on whether a model can call a tool correctly rather than how the visible tool menu shapes re…

An Assessment of Model-On-Model Deception

2024-05-10 · Julius Heitkoetter, Michael Gerovitch, Laker Newhouse

The trustworthiness of highly capable language models is put at risk when they are able to produce deceptive outputs. Moreover, when models are vulnerable to deception it undermines reliability. In this paper, we introdu…

MMLUmodel

Artificial intelligence and pediatrics: A synthetic mini review

2018-02-16 · Peter Kokol, Jernej Završnik, Helena Blažun Vošner

The use of artificial intelligence intelligencein medicine can be traced back to 1968 when Paycha published his paper Le diagnostic a l'aide d'intelligences artificielle, presentation de la premiere machine diagnostri. F…

Decision MakingDiagnostic