paper-with-me

Papers

Towards Reliable Dermatology Evaluation Benchmarks

2023-09-13 · Fabian Gröger, Simone Lionetti, Philippe Gottfrois, Alvaro Gonzalez-Jimenez, Matthew Groh, Roxana Daneshjou, Labelling Consortium, Alexander A. Navarini, Marc Pouly

Benchmark datasets for digital dermatology unwittingly contain inaccuracies that reduce trust in model performance estimates. We propose a resource-efficient data-cleaning protocol to identify issues that escaped previous curation. The protocol leverages an existing algorithmic cleaning strategy and is followed by a confirmation process terminated by an intuitive stopping criterion. Based on confirmation by multiple dermatologists, we remove irrelevant samples and near duplicates and estimate the percentage of label errors in six dermatology image datasets for model evaluation promoted by the International Skin Imaging Collaboration. Along with this paper, we publish revised file lists for each dataset which should be used for model evaluation. Our work paves the way for more trustworthy performance assessment in digital dermatology.

📄 PDF Abstract BibTeX arXiv:2309.06961

Code (2)

digital-dermatology/selfclean-revised-benchmarks 공식 구현
digital-dermatology/selfclean pytorch

Similar Papers 제목 키워드 기반

Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives

2025-11-12 · Yuhao Shen, Jiahe Qian, Shuping Zhang, Zhangtianyi Chen 외 arxiv

Multimodal large language models (LLMs) are increasingly used to generate dermatology diagnostic narratives directly from images. However, reliable evaluation remains the primary bottleneck for responsible clinical deplo…

Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology

2026-05-01 · Roy Jiang, Hyunjae Kim, Zhenyue Qin, Morten Lee 외 arxiv

Multimodal large language models (MLLMs) have demonstrated promise on publicly available dermatology benchmarks. However, benchmark performance may not generalize to real-world dermatologic decision-making. To quantify t…

Disparities in Dermatology AI: Assessments Using Diverse Clinical Images

2021-11-15 · Roxana Daneshjou, Kailas Vodrahalli, Weixin Liang, Roberto A Novoa 외

More than 3 billion people lack access to care for skin disease. AI diagnostic tools may aid in early skin cancer detection; however most models have not been assessed on images of diverse skin tones or uncommon diseases…

Diagnostic

DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning

2026-01-20 · Abdurrahim Yilmaz, Ozan Erdem, Ece Gokyayla, Ayda Acar 외 arxiv

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesi…

Visual Question Answering

Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology

2026-08-17 · Fabian Gröger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti 외 arxiv

Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing. These test-time images are ordinary clinical…