paper-with-me

홈 › Papers

Beyond Accuracy: Assessing Calibration of Geospatial Foundation Models and Their Sensitivity to Distribution Shifts

2026-08-17 · Nils Lehmann, Jakob Gawlikowski, Burak Ekim, Isaac Corley, Xiao Xiang Zhu arxiv

Geospatial Foundation Models (GeoFMs) are most commonly ranked and selected by accuracy on standard benchmark conditions via averaged ranks. We show that this protocol is too narrow: the promised deployment in critical EO tasks requires further angles of analysis, mainly calibration, the agreement between a model's confidence and its correctness. Across 16 frozen encoders, four classification and five segmentation datasets, and two orthogonal stress axes, every encoder degrades as corruption intensifies, and the ranking changes as well. Across the four classification benchmarks, EO-pretrained and ImageNet-pretrained encoders are indistinguishable on clean accuracy and clean calibration, and EO pretraining provides no more stability under shift than ImageNet pretraining. Under shift the GeoFMs drift further into overconfidence than the ImageNet-pretrained encoders, at every grade and in every corruption family. A centered kernel alignment (CKA) analysis ties this to representational rigidity: EO-pretrained embeddings move less under corruption while losing just as much task information and remaining overconfident. We apply three commonly explored uncertainty quantification methods and find that temperature scaling and deep ensembles cannot counteract the degradation, while a Gaussian-process probe roughly halves ECE under severe cloud only by tripling it on clean data. In selective prediction experiments, we find that confidence-based abstention cannot defer around confidently wrong predictions, and advocate that benchmark rankings and evaluations should therefore operate across a multitude of conditions and metrics to more holistically evaluate model development progress and close the gap to real world deployment scenarios.

📄 PDF Abstract BibTeX arXiv:2608.16614

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?

2025-10-17 · Coen Adler, Yuxin Chang, Felix Draxler, Samar Abdi 외 arxiv

The recent development of foundation models for time series data has generated considerable interest in using such models across a variety of applications. Although foundation models achieve state-of-the-art predictive p…

Scalable Geospatial Data Generation Using AlphaEarth Foundations Model

2025-08-15 · Luc Houriez, Sebastian Pilarski, Behzad Vahedi, Ali Ahmadalipour 외 arxiv

High-quality labeled geospatial datasets are essential for extracting insights and understanding our planet. Unfortunately, these datasets often do not span the entire globe and are limited to certain geographic regions …

Trustworthiness Calibration Framework for Phishing Email Detection Using Large Language Models

2025-11-06 · Daniyal Ganiuly, Assel Smaiyl arxiv

Phishing emails continue to pose a persistent challenge to online communication, exploiting human trust and evading automated filters through realistic language and adaptive tactics. While large language models (LLMs) su…

Text Classification

Segment Anything Model Can Not Segment Anything: Assessing AI Foundation Model's Generalizability in Permafrost Mapping

2024-01-16 · Wenwen Li, Chia-Yu Hsu, Sizhe Wang, Yezhou Yang 외

This paper assesses trending AI foundation models, especially emerging computer vision foundation models and their performance in natural landscape feature segmentation. While the term foundation model has quickly garner…

Instance SegmentationSemantic Segmentation

Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization

2026-01-12 · Thomas Snyder, H. Lexie Yang, Stefan Schnake, Steffen Schotthöfer arxiv

Deploying geospatial foundation models on resource-constrained edge devices demands compact architectures that maintain high downstream performance. However, their large parameter counts and the accuracy loss often induc…

Transfer Learning