paper-with-me

Papers

The Normalized Maximum Likelihood for Regular Non-Smooth Models: Measure-Theoretic Foundations and Geometric Sampling

2026-05-23 · Trenton Lau, Gary P. T. Choi arxiv

The Normalized Maximum Likelihood (NML) codelength, or stochastic complexity, represents a principled criterion for universal coding. While recent coarea-based formulations provided a calculation method for smooth models, this framework collapses for the non-smooth estimators ubiquitous in modern machine learning (e.g., Lasso, Sparse SVMs). In this work, we provide a rigorous framework for computing the NML for regular path-differentiable Lipschitz (PDL) estimators. By applying classical geometric measure theory and bridging the coarea formula with conservative Jacobians, we prove that the stochastic complexity for non-smooth models is well-posed and theoretically consistent with the outputs of modern Automatic Differentiation. To compute this quantity exactly, we introduce the Propose-and-Project Metropolis-Hastings (PDL-PPMH) sampler, a geometric MCMC algorithm capable of traversing the non-differentiable level sets of the maximum likelihood estimator. We theoretically justify its components, including a stochastic tangent space proposal and a provably convergent non-smooth projection solver. We demonstrate the method's robustness by sampling from a high-dimensional Lasso posterior ($P=2000$), while simultaneously quantifying the computational scaling that governs the trade-off between exactness and mixing time. Crucially, we empirically demonstrate that our exact NML criterion provides a highly data-efficient alternative to cross-validation, achieving statistically indistinguishable predictive optima without requiring data splitting. Altogether, our work paves the way for the theoretical analysis of the NML codelength for regular non-smooth models.

📄 PDF Abstract BibTeX arXiv:2605.24477

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sub-Gaussian Concentration and Entropic Normality of the Maximum Likelihood Estimator

2026-05-08 · Leighton P. Barnes, Alex Dytso arxiv

It is well known that, under standard regularity conditions, the maximum likelihood estimator (MLE) satisfies a central limit theorem and converges in distribution to a Gaussian random variable as the sample size grows. …

Quotient Normalized Maximum Likelihood Criterion for Learning Bayesian Network Structures

2024-08-27 · Tomi Silander, Janne Leppä-aho, Elias Jääsaari, Teemu Roos

We introduce an information theoretic criterion for Bayesian network structure learning which we call quotient normalized maximum likelihood (qNML). In contrast to the closely related factorized normalized maximum likeli…

Deep pNML: Predictive Normalized Maximum Likelihood for Deep Neural Networks

2019-04-28 · Koby Bibas, Yaniv Fogel, Meir Feder

The Predictive Normalized Maximum Likelihood (pNML) scheme has been recently suggested for universal learning in the individual setting, where both the training and test samples are individual data. The goal of universal…

Maximum Entropy Vector Kernels for MIMO system identification

2015-08-12 · Giulia Prando, Gianluigi Pillonetto, Alessandro Chiuso

Recent contributions have framed linear system identification as a nonparametric regularized inverse problem. Relying on $\ell_2$-type regularization which accounts for the stability and smoothness of the impulse respons…

Bayesian Properties of Normalized Maximum Likelihood and its Fast Computation

2014-01-28 · Andrew Barron, Teemu Roos, Kazuho Watanabe

The normalized maximized likelihood (NML) provides the minimax regret solution in universal data compression, gambling, and prediction, and it plays an essential role in the minimum description length (MDL) method of sta…

Data CompressionPrediction