paper-with-me

Papers

Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030

2025-05-12 · Mouxiao Bian, Rongzhao Zhang, Chao Ding, Xinwei Peng, Jie Xu

Large Language Models (LLMs) are poised to transform healthcare under China's Healthy China 2030 initiative, yet they introduce new ethical and patient-safety challenges. We present a novel 12,000-item Q&A benchmark covering 11 ethics and 9 safety dimensions in medical contexts, to quantitatively evaluate these risks. Using this dataset, we assess state-of-the-art Chinese medical LLMs (e.g., Qwen 2.5-32B, DeepSeek), revealing moderate baseline performance (accuracy 42.7% for Qwen 2.5-32B) and significant improvements after fine-tuning on our data (up to 50.8% accuracy). Results show notable gaps in LLM decision-making on ethics and safety scenarios, reflecting insufficient institutional oversight. We then identify systemic governance shortfalls-including the lack of fine-grained ethical audit protocols, slow adaptation by hospital IRBs, and insufficient evaluation tools-that currently hinder safe LLM deployment. Finally, we propose a practical governance framework for healthcare institutions (embedding LLM auditing teams, enacting data ethics guidelines, and implementing safety simulation pipelines) to proactively manage LLM risks. Our study highlights the urgent need for robust LLM governance in Chinese healthcare, aligning AI innovation with patient safety and ethical standards.

📄 PDF Abstract BibTeX arXiv:2505.07205

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingEthics

Similar Papers 제목 키워드 기반

A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare

2025-02-21 · Manar Aljohani, Jun Hou, Sindhura Kommu, Xuan Wang

The application of large language models (LLMs) in healthcare has the potential to revolutionize clinical decision-making, medical research, and patient care. As LLMs are increasingly integrated into healthcare systems, …

Decision MakingFairnessMisinformation

LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs

2024-10-18 · Yujun Zhou, Jingdong Yang, Kehan Guo, Pin-Yu Chen 외

Laboratory accidents pose significant risks to human life and property, underscoring the importance of robust safety protocols. Despite advancements in safety training, laboratory personnel may still unknowingly engage i…

BenchmarkingFairnessMultiple-choice

Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data

2025-05-15 · Adel ElZemity, Budi Arief, Shujun Li

The integration of large language models (LLMs) into cyber security applications presents significant opportunities, such as enhancing threat analysis and malware detection, but can also introduce critical risks and safe…

Malware DetectionSafety Alignment

When LLM Therapists Become Salespeople: Evaluating Large Language Models for Ethical Motivational Interviewing

2025-03-30 · Haein Kong, Seonghyeon Moon

Large language models (LLMs) have been actively applied in the mental health field. Recent research shows the promise of LLMs in applying psychotherapy, especially motivational interviewing (MI). However, there is a lack…

EthicsResponse Generation

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

2025-07-22 · Zheng Hui, Yijiang River Dong, Ehsan Shareghi, Nigel Collier arxiv

As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their domain-specific safety and compliance becomes critical. While prior work …