paper-with-me

홈 › Papers

Clean & Clear: Feasibility of Safe LLM Clinical Guidance

2025-03-26 · Julia Ive, Felix Jozsa, Nick Jackson, Paulina Bondaronek, Ciaran Scott Hill, Richard Dobson

Background: Clinical guidelines are central to safe evidence-based medicine in modern healthcare, providing diagnostic criteria, treatment options and monitoring advice for a wide range of illnesses. LLM-empowered chatbots have shown great promise in Healthcare Q&A tasks, offering the potential to provide quick and accurate responses to medical inquiries. Our main objective was the development and preliminary assessment of an LLM-empowered chatbot software capable of reliably answering clinical guideline questions using University College London Hospital (UCLH) clinical guidelines. Methods: We used the open-weight Llama-3.1-8B LLM to extract relevant information from the UCLH guidelines to answer questions. Our approach highlights the safety and reliability of referencing information over its interpretation and response generation. Seven doctors from the ward assessed the chatbot's performance by comparing its answers to the gold standard. Results: Our chatbot demonstrates promising performance in terms of relevance, with ~73% of its responses rated as very relevant, showcasing a strong understanding of the clinical context. Importantly, our chatbot achieves a recall of 1.00 for extracted guideline lines, substantially minimising the risk of missing critical information. Approximately 78% of responses were rated satisfactory in terms of completeness. A small portion (~14.5%) contained minor unnecessary information, indicating occasional lapses in precision. The chatbot' showed high efficiency, with an average completion time of 10 seconds, compared to 30 seconds for human respondents. Evaluation of clinical reasoning showed that 72% of the chatbot's responses were without flaws. Our chatbot demonstrates significant potential to speed up and improve the process of accessing locally relevant clinical information for healthcare professionals.

📄 PDF Abstract BibTeX arXiv:2503.20953

Code (0)

등록된 구현이 없습니다.

Tasks

ChatbotDiagnosticResponse Generation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Translating Milli/Microrobots with A Value-Centered Readiness Framework

2025-10-14 · Hakan Ceylan, Edoardo Sinibaldi, Sanjay Misra, Pankaj J. Pasricha 외 arxiv

Untethered mobile milli/microrobots hold transformative potential for interventional medicine by enabling more precise and entirely non-invasive diagnosis and therapy. Realizing this promise requires bridging the gap bet…

SafeCFG: Redirecting Harmful Classifier-Free Guidance for Safe Generation

2024-12-20 · Jiadong Pan, Hongcheng Gao, Liang Li, Zheng-Jun Zha 외

Diffusion models (DMs) have demonstrated exceptional performance in text-to-image (T2I) tasks, leading to their widespread use. With the introduction of classifier-free guidance (CFG), the quality of images generated by …

Image Generation

Towards automated patient data cleaning using deep learning: A feasibility study on the standardization of organ labeling

2017-12-30 · Timothy Rozario, Troy Long, Mingli Chen, Weiguo Lu 외

Data cleaning consumes about 80% of the time spent on data analysis for clinical research projects. This is a much bigger problem in the era of big data and machine learning in the field of medicine where large volumes o…

Safety and accuracy follow different scaling laws in clinical large language models

2026-05-05 · Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa, Mahshad Lotfinia 외 arxiv

Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation that higher accuracy implies safer behavior. This assumption is inco…

In vivo feasibility study of humanoid robots in surgery

2026-07-08 · Zekai Liang, Nikita Thareja, Peihan Zhang, Calvin Joyce 외 arxiv

Recent advances in actuation, control and learning have rapidly pushed humanoid robots from a distant vision towards near-term real-world deployment. Healthcare is a particularly pressing domain, in which staffing shorta…