paper-with-me

Papers

Discovery of Hidden Miscalibration Regimes

2026-05-13 · Katarzyna Kobalczyk, Mihaela van der Schaar arxiv

Calibration is commonly evaluated by comparing model confidence with its empirical correctness, implicitly treating reliability as a function of the confidence score alone. However, this view can hide substantial structure: models may be systematically overconfident on some kinds of inputs and underconfident on others, causing global reliability diagnostics to obscure localised calibration failures. To address this, we formulate the problem of discovering hidden miscalibration regimes without assuming access to predefined data slices. We define the corresponding miscalibration field and propose a diagnostic framework for estimating it. Our approach learns a calibration-aware representation of the input space and estimates signed local miscalibration by kernel smoothing in the learned geometry. Across four real-world LLM benchmarks and twelve LLMs, we find that input-dependent calibration heterogeneity is prevalent. We further show that the discovered fields are actionable: they support local confidence correction and reduce calibration error in systematically miscalibrated regions where confidence-based methods such as isotonic regression and temperature scaling are less effective.

📄 PDF Abstract BibTeX arXiv:2605.13484

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IHCV: Discovery of Hidden Time-Dependent Control Variables in Non-Linear Dynamical Systems

2023-04-05 · Juan Munoz, Subash Balsamy, Juan P. Bernal-Tamayo, Ali Balubaid 외

Discovering non-linear dynamical models from data is at the core of science. Recent progress hinges upon sparse regression of observables using extensive libraries of candidate functions. However, it remains challenging …

Benchmarking

Empirically Calibrated Conditional Independence Tests

2026-02-24 · Milleno Pan, Antoine de Mathelin, Wesley Tansey arxiv

Conditional independence tests (CIT) are widely used for causal discovery and feature selection. Even with false discovery rate (FDR) control procedures, they often fail to provide frequentist guarantees in practice. We …

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

2026-08-26 · Tim Schopf, Tobias Schreieder, Akiko Aizawa arxiv

Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we invest…

Identifiable Markov Switching Models with Instantaneous Effects and Exponential Families

2026-06-01 · Roel Hulsman, Carles Balsells-Rodas, Sara Magliacane arxiv

Temporal systems often exhibit non-stationary behaviour, such as seasonal climate variation or glucose fluctuations in patients with type-1 diabetes. One way to model non-stationarity is through discrete latent regimes, …

False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs

2025-10-16 · Akira Okutomi arxiv

High-confidence errors in large language models are often treated as fragile failures. We study an alternative: some errors may be false fixed points, locally stable, internally coherent, and confidently wrong. This sepa…