paper-with-me

Papers

DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning

2026-01-20 · Abdurrahim Yilmaz, Ozan Erdem, Ece Gokyayla, Ayda Acar, Burc Bugra Dagtas, Dilara Ilhan Erdil, Gulsum Gencoglan, Burak Temelkuran arxiv

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition. While valuable for recognition, such datasets cannot assess the full visual understanding, language grounding, and clinical reasoning capabilities of multimodal models. Visual question answering (VQA) benchmarks are required to evaluate how models interpret dermatological images, reason over fine-grained morphology, and generate clinically meaningful descriptions. We introduce DermaBench, a clinician-annotated dermatology VQA benchmark built on the Diverse Dermatology Images (DDI) dataset. DermaBench comprises 656 clinical images from 570 unique patients spanning Fitzpatrick skin types I-VI. Using a hierarchical annotation schema with 22 main questions (single-choice, multi-choice, and open-ended), expert dermatologists annotated each image for diagnosis, anatomic site, lesion morphology, distribution, surface features, color, and image quality, together with open-ended narrative descriptions and summaries, yielding approximately 14.474 VQA-style annotations. DermaBench is released as a metadata-only dataset to respect upstream licensing and is publicly available at Harvard Dataverse.

📄 PDF Abstract BibTeX arXiv:2601.14084

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

SkinCon: A skin disease dataset densely annotated by domain experts for fine-grained model debugging and analysis

2023-02-01 · Roxana Daneshjou, Mert Yuksekgonul, Zhuo Ran Cai, Roberto Novoa 외

For the deployment of artificial intelligence (AI) in high-risk settings, such as healthcare, methods that provide interpretability/explainability or allow fine-grained error analysis are critical. Many recent methods fo…

Interpretable Machine Learning

SkinCAP: A Multi-modal Dermatology Dataset Annotated with Rich Medical Captions

2024-05-28 · Juexiao Zhou, Liyuan Sun, Yan Xu, wenbin liu 외

With the widespread application of artificial intelligence (AI), particularly deep learning (DL) and vision-based large language models (VLLMs), in skin disease diagnosis, the need for interpretability becomes crucial. H…

Enhanced Dermatology Image Quality Assessment via Cross-Domain Training

2025-06-19 · Ignacio Hernández Montilla, Alfonso Medela, Paola Pasquali, Andy Aguilar 외

Teledermatology has become a widely accepted communication method in daily clinical practice, enabling remote care while showing strong agreement with in-person visits. Poor image quality remains an unsolved problem in t…

Image Quality Assessment

TrueImage: A Machine Learning Algorithm to Improve the Quality of Telehealth Photos

2020-10-01 · Kailas Vodrahalli, Roxana Daneshjou, Roberto A Novoa, Albert Chiou 외

Telehealth is an increasingly critical component of the health care ecosystem, especially due to the COVID-19 pandemic. Rapid adoption of telehealth has exposed limitations in the existing infrastructure. In this paper, …

BIG-bench Machine Learning

Towards Scalable Foundation Models for Digital Dermatology

2024-11-08 · Fabian Gröger, Philippe Gottfrois, Ludovic Amruthalingam, Alvaro Gonzalez-Jimenez 외

The growing demand for accurate and equitable AI models in digital dermatology faces a significant challenge: the lack of diverse, high-quality labeled data. In this work, we investigate the potential of domain-specific …

DiagnosticSelf-Supervised Learning