paper-with-me

홈 › Papers

Towards medical AI misalignment: a preliminary study

2025-05-22 · Barbara Puccio, Federico Castagna, Allan Tucker, Pierangelo Veltri

Despite their staggering capabilities as assistant tools, often exceeding human performances, Large Language Models (LLMs) are still prone to jailbreak attempts from malevolent users. Although red teaming practices have already identified and helped to address several such jailbreak techniques, one particular sturdy approach involving role-playing (which we named `Goofy Game') seems effective against most of the current LLMs safeguards. This can result in the provision of unsafe content, which, although not harmful per se, might lead to dangerous consequences if delivered in a setting such as the medical domain. In this preliminary and exploratory study, we provide an initial analysis of how, even without technical knowledge of the internal architecture and parameters of generative AI models, a malicious user could construct a role-playing prompt capable of coercing an LLM into producing incorrect (and potentially harmful) clinical suggestions. We aim to illustrate a specific vulnerability scenario, providing insights that can support future advancements in the field.

📄 PDF Abstract BibTeX arXiv:2505.18212

Code (0)

등록된 구현이 없습니다.

Tasks

Red Teaming

Similar Papers 제목 키워드 기반

Emergent Misalignment is Easy, Narrow Misalignment is Hard

2026-02-08 · Anna Soligo, Edward Turner, Senthooran Rajamanoharan, Neel Nanda arxiv

Finetuning large language models on narrowly harmful datasets can cause them to become emergently misaligned, giving stereotypically `evil' responses across diverse unrelated settings. Concerningly, a pre-registered surv…

Right Prediction, Wrong Reasoning: Uncovering LLM Misalignment in RA Disease Diagnosis

2025-04-09 · Umakanta Maharana, Sarthak Verma, Avarna Agarwal, Prakashini Mruthyunjaya 외

Large language models (LLMs) offer a promising pre-screening tool, improving early disease detection and providing enhanced healthcare access for underprivileged communities. The early diagnosis of various diseases conti…

DiagnosticDisease PredictionPrediction

A Preliminary Approach for Learning Relational Policies for the Management of Critically Ill Children

2020-01-13 · Michael A. Skinner, Lakshmi Raman, Neel Shah, Abdelaziz Farhat 외

The increased use of electronic health records has made possible the automated extraction of medical policies from patient records to aid in the development of clinical decision support systems. We adapted a boosted Stat…

ManagementRelational ReasoningRespiratory Failure

Deformation-aware GAN for Medical Image Synthesis with Substantially Misaligned Pairs

2024-08-18 · Bowen Xin, Tony Young, Claire E Wainwright, Tamara Blake 외

Medical image synthesis generates additional imaging modalities that are costly, invasive or harmful to acquire, which helps to facilitate the clinical workflow. When training pairs are substantially misaligned (e.g., lu…

Image Generation

The Impact of Image Resolution on Biomedical Multimodal Large Language Models

2025-10-21 · Liangyu Chen, James Burgess, Jeffrey J Nirschl, Orr Zohar 외 arxiv

Imaging technologies are fundamental to biomedical research and modern medicine, requiring analysis of high-resolution images across various modalities. While multimodal large language models (MLLMs) show promise for bio…