paper-with-me

홈 › Papers

Hearsay: Vision-Language Medical Diagnoses Without an Image

2026-07-29 · Siddharth Vohra arxiv

When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads "suspected, based on demographics and classic pattern.'' GPT-5.4's effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings sharpen the problem. A hedged regime appears in which the prose acknowledges the missing image while the structured diagnosis field nevertheless names a disease, a dissociation invisible to prose-only audits. And Claude's dermatology effect collapses entirely when 'skin mole' is swapped for 'skin lesion' while GPT-5.4's is preserved, indicating that mirage is a family of distinct failure modes rather than a single phenomenon. Trustworthy VLM deployment in clinical pipelines requires auditing the structured output channel directly, and probe-word sensitivity should be treated as a first-class evaluation dimension

📄 PDF Abstract BibTeX arXiv:2607.26886

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HearSay Benchmark: Do Audio LLMs Leak What They Hear?

2026-01-07 · Jin Wang, Liang Lin, Kaiwen Luo, Weiliu Wang 외 arxiv

While Audio Large Language Models (ALLMs) have achieved remarkable progress in understanding and generation, their potential privacy implications remain largely unexplored. This paper takes the first step to investigate …

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI

2024-11-21 · Tianbin Li, Yanzhou Su, Wei Li, Bin Fu 외

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset cr…

Decision MakingLanguage ModelingLanguage ModellingQuestion Answering+1

Human-AI collectives produce the most accurate differential diagnoses

2024-06-21 · N. Zöller, J. Berger, I. Lin, N. Fu 외

Artificial intelligence systems, particularly large language models (LLMs), are increasingly being employed in high-stakes decisions that impact both individuals and society at large, often without adequate safeguards to…

Common Sense Reasoning

Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks

2025-11-10 · Lingran Song, Yucheng Zhou, Jianbing Shen arxiv

Despite significant progress in pixel-level medical image analysis, existing medical image segmentation models rarely explore medical segmentation and diagnosis tasks jointly. However, it is crucial for patients that mod…

Medical Image SegmentationMedical Diagnosis

Training Medical Large Vision-Language Models with Abnormal-Aware Feedback

2025-01-02 · Yucheng Zhou, Lingran Song, Jianbing Shen

Existing Medical Large Vision-Language Models (Med-LVLMs), which encapsulate extensive medical knowledge, demonstrate excellent capabilities in understanding medical images and responding to human queries based on these …

Anomaly DetectionVisual Localization