paper-with-me

홈 › Papers

How Good is my Histopathology Vision-Language Foundation Model? A Holistic Benchmark

2025-03-17 · Roba Al Majzoub, Hashmat Malik, Muzammal Naseer, Zaigham Zaheer, Tariq Mahmood, Salman Khan, Fahad Khan

Recently, histopathology vision-language foundation models (VLMs) have gained popularity due to their enhanced performance and generalizability across different downstream tasks. However, most existing histopathology benchmarks are either unimodal or limited in terms of diversity of clinical tasks, organs, and acquisition instruments, as well as their partial availability to the public due to patient data privacy. As a consequence, there is a lack of comprehensive evaluation of existing histopathology VLMs on a unified benchmark setting that better reflects a wide range of clinical scenarios. To address this gap, we introduce HistoVL, a fully open-source comprehensive benchmark comprising images acquired using up to 11 various acquisition tools that are paired with specifically crafted captions by incorporating class names and diverse pathology descriptions. Our Histo-VL includes 26 organs, 31 cancer types, and a wide variety of tissue obtained from 14 heterogeneous patient cohorts, totaling more than 5 million patches obtained from over 41K WSIs viewed under various magnification levels. We systematically evaluate existing histopathology VLMs on Histo-VL to simulate diverse tasks performed by experts in real-world clinical scenarios. Our analysis reveals interesting findings, including large sensitivity of most existing histopathology VLMs to textual changes with a drop in balanced accuracy of up to 25% in tasks such as Metastasis detection, low robustness to adversarial attacks, as well as improper calibration of models evident through high ECE values and low model prediction confidence, all of which can affect their clinical implementation.

📄 PDF Abstract BibTeX arXiv:2503.12990

Code (1)

musk007/Histopathology_Benchmark 공식 구현 pytorch

Similar Papers 제목 키워드 기반

FLAVA: A Foundational Language And Vision Alignment Model

2021-12-08 · CVPR 2022 1 · Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 외

State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal…

Image RetrievalImage-to-Text RetrievalVisual ReasoningZero-shot Image Retrieval+2

Mind the Gap: Evaluating Patch Embeddings from General-Purpose and Histopathology Foundation Models for Cell Segmentation and Classification

2025-02-04 · Valentina Vadori, Antonella Peruffo, Jean-Marie Graïc, Livio Finos 외

Recent advancements in foundation models have transformed computer vision, driving significant performance improvements across diverse domains, including digital histopathology. However, the advantages of domain-specific…

Cell SegmentationDecoderInstance SegmentationModel Selection+3

Automated Histopathology Report Generation via Pyramidal Feature Extraction and the UNI Foundation Model

2026-02-18 · Ahmet Halici, Ece Tugba Cebeci, Musa Balci, Mustafa Cini 외 arxiv

Generating diagnostic text from histopathology whole slide images (WSIs) is challenging due to the gigapixel scale of the input and the requirement for precise, domain specific language. We propose a hierarchical vision …

Benchmarking Computational Pathology Foundation Models For Semantic Segmentation

2026-02-21 · Lavish Ramchandani, Aashay Tinaikar, Dev Kumar Das, Rohit Garg 외 arxiv

In recent years, foundation models such as CLIP, DINO,and CONCH have demonstrated remarkable domain generalization and unsupervised feature extraction capabilities across diverse imaging tasks. However, systematic and in…

Semantic SegmentationDomain GeneralizationNuclear Segmentation

Towards a Visual-Language Foundation Model for Computational Pathology

2023-07-24 · Ming Y. Lu, Bowen Chen, Drew F. K. Williamson, Richard J. Chen 외

The accelerated adoption of digital pathology and advances in deep learning have enabled the development of powerful models for various pathology tasks across a diverse array of diseases and patient cohorts. However, mod…

Contrastive Learningimage-classificationImage ClassificationImage to text+3