paper-with-me

홈 › Papers

DALPHIN: Benchmarking Digital Pathology AI Copilots Against Pathologists on an Open Multicentric Dataset

2026-05-05 · Carlijn Lems, Sander Moonemans, Natálie Klubíčková, Biagio Brattoli, Taebum Lee, Seokhwi Kim, Veronica Vilaplana, Laura Pons, Sapir Hochman, Mauricio Eduardo Suárez-Franck, Pedro Luis Fernandez, Julius Drachneris, Donatas Petroska, Renaldas Augulis, Arvydas Laurinavicius, Domingos Oliveira, Diana Montezuma, Anouk B. Bouwmeester, Dominique van Midden, Anne-Marie Vos, Shoko Vos, Jolique van Ipenburg, Maschenka Balkenhol, Koen Winkler, Iris Nagtegaal, Konnie Hebeda, Uta Flucke, Katrien Grünberg, Josef Skopal, Brinder S. Chohan, Jordi Temprana-Salvador, Enrico Munari, Luca Cima, Giulia Querzoli, Yosamin Gonzalez Belisario, Jaeike W. Faber, Geert J. L. H. van Leenders, Jan H. von der Thüsen, Lodewijk A. A. Brosens, Ronald R. de Krijger, Pieter Wesseling, Sandrine Florquin, Mateusz Maniewski, Adam Kowalewski, Robert Barna, Dina Tiniakos, Joan Lop Gros, Rogier Donders, Jake S. F. Maurits, Ming Yang Lu, Chengkuan Chen, Faisal Mahmood, Jeroen van der Laak, Nadieh Khalili, Frédérique Meeuwsen, Francesco Ciompi arxiv

Foundation models with visual question answering capabilities for digital pathology are emerging. Such unprecedented technology requires independent benchmarking to assess its potential in assisting pathologists in routine diagnostics. We created DALPHIN, the first multicentric open benchmark for pathology AI copilots, comprising 1236 images from 300 cases, spanning 130 rare to common diagnoses, 6 countries, and 14 subspecialties. The DALPHIN design and dataset are introduced alongside a human performance benchmark of 31 pathologists from 10 countries with varying expertise. We report results for two general-purpose (GPT-5, Gemini 2.5 Pro) and one pathology-specific copilot (PathChat+) for sequential and independent answer generation. We observed no statistically significant difference from expert-level performance in four of six tasks for PathChat, 2/6 tasks for Gemini, and 1/6 tasks for GPT. DALPHIN is publicly released with sequestered, indirectly accessible ground truth to foster robust and enduring benchmarking. Data, methods, and the evaluation platform are accessible through dalphin.grand-challenge.org.

📄 PDF Abstract BibTeX arXiv:2605.03544

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringAnswer Generation

Similar Papers 제목 키워드 기반

Improved statistical benchmarking of digital pathology models using pairwise frames evaluation

2023-06-07 · Ylaine Gerardin, John Shamshoian, Judy Shen, Nhat Le 외

Nested pairwise frames is a method for relative benchmarking of cell or tissue digital pathology models against manual pathologist annotations on a set of sampled patches. At a high level, the method compares agreement b…

BenchmarkingClassification

TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots

2025-11-20 · Tianyu Liu, Weihao Xuan, Hao Wu, Peter Humphrey 외 arxiv

Advances in AI have introduced several strong models in computational pathology to usher it into the era of multi-modal diagnosis, analysis, and interpretation. However, the current pathology-specific visual language mod…

Reinforcement Learning

Quantitative Benchmarking of Anomaly Detection Methods in Digital Pathology

2025-06-24 · Can Cui, Xindong Zheng, Ruining Deng, Quan Liu 외

Anomaly detection has been widely studied in the context of industrial defect inspection, with numerous methods developed to tackle a range of challenges. In digital pathology, anomaly detection holds significant potenti…

Anomaly DetectionArtifact DetectionBenchmarking

PhenoBench: A Comprehensive Benchmark for Cell Phenotyping

2025-07-04 · Claudia Winklmayr, Jerome Luescher, Nora Koreuber, Jannik Franzen 외 arxiv

Digital pathology has seen the advent of a wealth of foundational models (FM), yet to date their performance on cell phenotyping has not been benchmarked in a unified manner. We therefore propose PhenoBench: A comprehens…

Classification and Retrieval of Digital Pathology Scans: A New Dataset

2017-05-22 · Morteza Babaie, Shivam Kalra, Aditya Sriram, Christopher Mitcheltree 외

In this paper, we introduce a new dataset, \textbf{Kimia Path24}, for image classification and retrieval in digital pathology. We use the whole scan images of 24 different tissue textures to generate 1,325 test patches o…

BenchmarkingGeneral Classificationimage-classificationImage Classification+1