paper-with-me

홈 › Papers

ReXInTheWild: A Unified Benchmark for Medical Photograph Understanding

2026-03-19 · Oishi Banerjee, Sung Eun Kim, Alexandra N. Willauer, Julius M. Kernbach, Abeer Rihan Alomaish, Reema Abdulwahab S. Alghamdi, Hassan Rayhan Alomaish, Mohammed Baharoon, Xiaoman Zhang, Julian Nicolas Acosta, Christine Zhou, Pranav Rajpurkar arxiv

Everyday photographs taken with ordinary cameras are already widely used in telemedicine and other online health conversations, yet no comprehensive benchmark evaluates whether vision-language models can interpret their medical content. Analyzing these images requires both fine-grained natural image understanding and domain-specific medical reasoning, a combination that challenges both general-purpose and specialized models. We introduce ReXInTheWild, a benchmark of 955 clinician-verified multiple-choice questions spanning seven clinical topics across 484 photographs sourced from the biomedical literature. When evaluated on ReXInTheWild, leading multimodal large language models show substantial performance variation: Gemini-3 achieves 78% accuracy, followed by Claude Opus 4.5 (72%) and GPT-5 (68%), while the medical specialist model MedGemma reaches only 37%. A systematic error analysis also reveals four categories of common errors, ranging from low-level geometric errors to high-level reasoning failures and requiring different mitigation strategies. ReXInTheWild provides a challenging, clinically grounded benchmark at the intersection of natural image understanding and medical reasoning. The dataset is available on HuggingFace.

📄 PDF Abstract BibTeX arXiv:2603.19517

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

2026-08-19 · Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang 외 arxiv

Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absenc…

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation

2025-10-09 · Kang Liao, Size Wu, Zhonghua Wu, Linyi Jin 외 arxiv

Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric multimodal model that extends spatial awa…

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

2025-10-17 · Junzhi Ning, Wei Li, Cheng Tang, Jiashi Lin 외 arxiv

Medical workflows routinely combine reading images with producing visual and textual outputs, making both image understanding and generation central to medical AI. Most existing systems, however, address these abilities …

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

2025-08-05 · Zhaoxu Li, Chenqi Kong, Yi Yu, Qiangqiang Wu 외 arxiv

Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previou…

Scene Understanding

DinoDental: Benchmarking DINOv3 as a Unified Vision Encoder for Dental Image Analysis

2026-03-30 · Kun Tang, Xinquan Yang, Mianjie Zheng, Xuefen Liu 외 arxiv

The scarcity and high cost of expert annotations in dental imaging present a significant challenge for the development of AI in dentistry. DINOv3, a state-of-the-art, self-supervised vision foundation model pre-trained o…

Instance Segmentation