paper-with-me

홈 › Papers

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?

2025-05-30 · Tianhong Zhou, Yin Xu, Yingtao Zhu, Chuxi Xiao, Haiyang Bian, Lei Wei, Xuegong Zhang

Vision-language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks do not systematically evaluate whether these models truly reason like human clinicians or merely imitate superficial patterns. To address this gap, we propose DrVD-Bench, the first multimodal benchmark for clinical visual reasoning. DrVD-Bench consists of three modules: Visual Evidence Comprehension, Reasoning Trajectory Assessment, and Report Generation Evaluation, comprising a total of 7,789 image-question pairs. Our benchmark covers 20 task types, 17 diagnostic categories, and five imaging modalities-CT, MRI, ultrasound, radiography, and pathology. DrVD-Bench is explicitly structured to reflect the clinical reasoning workflow from modality recognition to lesion identification and diagnosis. We benchmark 19 VLMs, including general-purpose and medical-specific, open-source and proprietary models, and observe that performance drops sharply as reasoning complexity increases. While some models begin to exhibit traces of human-like reasoning, they often still rely on shortcut correlations rather than grounded visual understanding. DrVD-Bench offers a rigorous and structured evaluation framework to guide the development of clinically trustworthy VLMs.

📄 PDF Abstract BibTeX arXiv:2505.24173

Code (1)

jerry-boss/drvd-bench 공식 구현

Tasks

DiagnosticMedical Image AnalysisVisual ReasoningZero-shot Generalization

Similar Papers 제목 키워드 기반

High Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression Algorithm

2025-03-01 · CVPR 2025 1 · Zhaoyi Tian, Feifeng Wang, Shiwei Wang, ZiHao Zhou 외

Recently, learned video compression (LVC) is undergoing a period of rapid development. However, due to absence of large and high-quality high dynamic range (HDR) video training data, LVC on HDR video is still unexplored.…

Video Compression

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?

2026-05-11 · Maximilian Triebel, Marco Menner, Dominik Helfenstein arxiv

Vision-Language(-Action) Models (VLMs) are increasingly applied to interactive environments, yet existing benchmarks often overlook the complex physical reasoning required for point-and-click puzzle games. This paper int…

Logical ReasoningVisual Grounding

MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models

2025-05-26 · Anh Thai, Stefan Stojanov, Zixuan Huang, Bikram Boote 외

This paper introduces MEBench, a novel benchmark for evaluating mutual exclusivity (ME) bias, a cognitive phenomenon observed in children during word learning. Unlike traditional ME tasks, MEBench further incorporates sp…

Spatial Reasoning

NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

2024-10-18 · Baiqi Li, Zhiqiu Lin, Wenxuan Peng, Jean de Dieu Nyandwi 외

Vision-language models (VLMs) have made significant progress in recent visual-question-answering (VQA) benchmarks that evaluate complex visio-linguistic reasoning. However, are these models truly effective? In this work,…

AttributeQuestion AnsweringTAGVisual Question Answering+1

Verbal Process Supervision Elicits Better Coding Agents

2025-03-24 · Hao-Yuan Chen, Cheng-Pong Huang, Jui-Ming Yao

The emergence of large language models and their applications as AI agents have significantly advanced state-of-the-art code generation benchmarks, transforming modern software engineering tasks. However, even with test-…

Code Generation