paper-with-me

홈 › Papers

The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models

2026-07-05 · Dhyey Yajnik, Amina Asif, Fayyaz Minhas arxiv

How robust and generalisable are pathology foundation models and have their scaling limites been reached? We benchmarked twelve pathology foundation models (PFMs) and ResNet baselines using our Robustness Evaluation and Enhancement Toolbox (REET) across eleven clinically realistic perturbations and a dissimilarity-driven Non-Redundant K-fold validation (NR-Kfold) protocol. We introduce a Perturbation Performance Index (PPI) to summarise accuracy trends under controlled perturbation sweeps and analyse robustness scaling with parameter count. We show that PFMs consistently outperform CNNs in both robustness and domain generalisation, yet model scaling shows diminishing returns: mid-sized models such (UNI2/Virchow-2 etc.) achieve comparable or greater resilience than larger systems. NR-Kfold analysis further reveals systematic accuracy loss and increased variability when training-test similarity is broken, underscoring the need for explicit distribution-shift evaluation. These findings suggest that the next generation of pathology foundation models must prioritise data quality, multimodality information and domain alignment over parameter count to achieve genuine clinical reliability.

📄 PDF Abstract BibTeX arXiv:2607.04401

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are nuclear masks all you need for improved out-of-domain generalisation? A closer look at cancer classification in histopathology

2024-11-14 · Dhananjay Tomar, Alexander Binder, Andreas Kleppe

Domain generalisation in computational histopathology is challenging because the images are substantially affected by differences among hospitals due to factors like fixation and staining of tissue and imaging equipment.…

AllCancer ClassificationData AugmentationNuclear Segmentation

Using Videos to Evaluate Image Model Robustness

2019-04-22 · Keren Gu, Brandon Yang, Jiquan Ngiam, Quoc Le 외

Human visual systems are robust to a wide range of image transformations that are challenging for artificial networks. We present the first study of image model robustness to the minute transformations found across video…

model

Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite Imagery

2025-03-24 · CVPR 2025 1 · Sara Al-Emadi, Yin Yang, Ferda Ofli

Object detectors have achieved remarkable performance in many applications; however, these deep learning models are typically designed under the i.i.d. assumption, meaning they are trained and evaluated on data sampled f…

BenchmarkingHumanitarianObjectobject-detection+1

PathBench-MIL: A Comprehensive AutoML and Benchmarking Framework for Multiple Instance Learning in Histopathology

2025-12-19 · Siemen Brussee, Pieter A. Valkema, Jurre A. J. Weijer, Thom Doeleman 외 arxiv

We introduce PathBench-MIL, an open-source AutoML and benchmarking framework for multiple instance learning (MIL) in histopathology. The system automates end-to-end MIL pipeline construction, including preprocessing, fea…

Multiple Instance Learning

What's in a Name? Are BERT Named Entity Representations just as Good for any other Name?

2020-07-14 · WS 2020 7 · Sriram Balasubramanian, Naman jain, Gaurav Jindal, Abhijeet Awasthi 외

We evaluate named entity representations of BERT-based NLP models by investigating their robustness to replacements from the same typed class in the input. We highlight that on several tasks while such perturbations are …