paper-with-me

Papers

Good Rankings, Wrong Probabilities: A Calibration Audit of Multimodal Cancer Survival Models

2026-04-05 · Sajad Ghawami arxiv

Multimodal deep learning models that fuse whole-slide histopathology images with genomic data have achieved strong discriminative performance for cancer survival prediction, as measured by the concordance index. Yet whether the survival probabilities derived from these models - either directly from native outputs or via standard post-hoc reconstruction - are calibrated remains largely unexamined. We conduct, to our knowledge, the first systematic fold-level 1-calibration audit of multimodal WSI-genomics survival architectures, evaluating native discrete-time survival outputs (Experiment A: 3 models on TCGA-BRCA) and Breslow-reconstructed survival curves from scalar risk scores (Experiment B: 11 architectures across 5 TCGA cancer types). In Experiment A, all three models fail 1-calibration on a majority of folds (12 of 15 fold-level tests reject after Benjamini-Hochberg correction). Across the full 290 fold-level tests, 166 reject the null of correct calibration at the median event time after Benjamini-Hochberg correction (FDR = 0.05). MCAT achieves C-index 0.817 on GBMLGG yet fails 1-calibration on all five folds. Gating-based fusion is associated with better calibration; bilinear and concatenation fusion are not. Post-hoc Platt scaling reduces miscalibration at the evaluated horizon (e.g., MCAT: 5/5 folds failing to 2/5) without affecting discrimination. The concordance index alone is insufficient for evaluating survival models intended for clinical use.

📄 PDF Abstract BibTeX arXiv:2604.04239

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Deep Learning

Similar Papers 제목 키워드 기반

Calibrated Preference Learning: The Case of Label Ranking

2026-05-28 · Santo M. A. R. Thies, Viktor Bengs, Timo Kaufmann, Sebastian J. Vollmer 외 arxiv

Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively studied for classification and regression, calibration has not been formal…

Truthful Calibration Errors for Multi-Class Prediction

2025-10-07 · Yuxuan Lu, Yifan Wu, Jason Hartline, Lunjia Hu arxiv

Calibrated predictions are useful because their numerical values can be interpreted as probabilities. Calibration errors are therefore widely used to evaluate, compare, and tune probabilistic predictors. Recently, Haghta…

Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model

2026-08-24 · Shashwat Pandey, Satwik Pandey, Suresh Raghu arxiv

Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device models now ship to hundreds of millions of devices with no server-side moderation, and the configuration developers can …

Single-Query Black-Box Calibration Auditing via Logit Bias

2026-09-04 · Roman Plaud, Antoine Saillenfest, Matthieu Labeau, Thomas Bonald 외 arxiv

Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the continuous output probabilities required by s…

Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only

2024-10-14 · Jihan Yao, Wenxuan Ding, Shangbin Feng, Lucy Lu Wang 외

In the absence of abundant reliable annotations for challenging tasks and contexts, how can we expand the frontier of LLM capabilities with potentially wrong answers? We focus on two research questions: (1) Can LLMs gene…