paper-with-me

홈 › Papers

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models

2026-05-20 · Sunday Oyinlola Ogundoyin, Muhammad Ikram, Rahat Masood arxiv

Medical large language models (LLMs), including custom medical GPTs (MedGPTs) and open-source models, are increasingly deployed on web platforms to provide clinical guidance. However, they pose risks of hallucination, policy noncompliance, and unsafe design. We conduct a large-scale assessment of 6,233 MedGPTs, evaluating a stratified sample of 1,500, together with 10 open-source LLMs. We introduce two frameworks: MedGPT-HEval for hallucination detection and an LLM-based pipeline for assessing policy violations and developer intent. Our results show that 25-30% of MedGPTs exhibit low factual accuracy, with bottom- and middle-tier models at highest risk; 33.6-54.3% violate operational thresholds, and 57.06% of Action-enabled models lack adequate privacy disclosures. Compared with open-source models, MedGPTs achieve higher factual accuracy and semantic alignment, though open-source models are more stable. These results reveal systemic gaps in hallucination and compliance, highlighting the need for multi-metric evaluation and stronger safeguards. We release HAA-MedGPT, a structured dataset that supports future research on the safety of web-facing medical LLMs.

📄 PDF Abstract BibTeX arXiv:2605.20591

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI Generated Child Sexual Abuse Material -- What's the Harm?

2025-10-03 · Caoilte Ó Ciardha, John Buckley, Rebecca S. Portnoff arxiv

The development of generative artificial intelligence (AI) tools capable of producing wholly or partially synthetic child sexual abuse material (AI CSAM) presents profound challenges for child protection, law enforcement…

Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse

2025-02-24 · Garrett Wilson, Geoffrey Goh, Yan Jiang, Ajay Gupta 외

Detecting phishing, spam, fake accounts, data scraping, and other malicious activity in online social networks (OSNs) is a problem that has been studied for well over a decade, with a number of important results. Nearly …

Abuse Detection

A Unified Taxonomy of Harmful Content

2020-11-01 · EMNLP (ALW) 2020 11 · Michele Banko, Brendon MacKeen, Laurie Ray

The ability to recognize harmful content within online communities has come into focus for researchers, engineers and policy makers seeking to protect users from abuse. While the number of datasets aiming to capture form…

A Just and Comprehensive Strategy for Using NLP to Address Online Abuse

2019-06-04 · ACL 2019 7 · David Jurgens, Eshwar Chandrasekharan, Libby Hemphill

Online abusive behavior affects millions and the NLP community has attempted to mitigate this problem by developing technologies to detect abuse. However, current methods have largely focused on a narrow definition of ab…

Position

AAA: Fair Evaluation for Abuse Detection Systems Wanted

2021-06-21 · ACM Web Science 2021 6 · Agostina Calabrese, Michele Bevilacqua, Björn Ross, Rocco Tripodi 외

User-generated web content is rife with abusive language that can harm others and discourage participation. Thus, a primary research aim is to develop abuse detection systems that can be used to alert and support human m…

Abuse DetectionAbusive LanguageHate Speech DetectionSelection bias