paper-with-me

홈 › Papers

Informative missingness and its implications in semi-supervised learning

2025-12-04 · Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan arxiv

Semi-supervised learning (SSL) constructs classifiers using both labelled and unlabelled data. It leverages information from labelled samples, whose acquisition is often costly or labour-intensive, together with unlabelled data to enhance prediction performance. This defines an incomplete-data problem, which statistically can be formulated within the likelihood framework for finite mixture models that can be fitted using the expectation-maximisation (EM) algorithm. Ideally, one would prefer a completely labelled sample, as one would anticipate that a labelled observation provides more information than an unlabelled one. However, when the mechanism governing label absence depends on the observed features or the class labels or both, the missingness indicators themselves contain useful information. In certain situations, the information gained from modelling the missing-label mechanism can even outweigh the loss due to missing labels, yielding a classifier with a smaller expected error than one based on a completely labelled sample analysed. This improvement arises particularly when class overlap is moderate, labelled data are sparse, and the missingness is informative. Modelling such informative missingness thus offers a coherent statistical framework that unifies likelihood-based inference with the behaviour of empirical SSL methods.

📄 PDF Abstract BibTeX arXiv:2512.04392

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ICYM2I: The illusion of multimodal informativeness under missingness

2025-05-22 · Young Sang Choi, Vincent Jeanselme, Pierre Elias, Shalmali Joshi

Multimodal learning is of continued interest in artificial intelligence-based applications, motivated by the potential information gain from combining different types of data. However, modalities collected and curated du…

Informativeness

Semi-Supervised Mixture Models under the Concept of Missing at Radom with Margin Confidence and Aranda Ordaz Function

2026-01-21 · Jinyang Liao, Ziyang Lyu arxiv

This paper presents a semi-supervised learning framework for Gaussian mixture modelling under a Missing at Random (MAR) mechanism. The method explicitly parameterizes the missingness mechanism by modelling the probabilit…

Time series cluster kernels to exploit informative missingness and incomplete label information

2019-07-10 · Karl Øyvind Mikalsen, Cristina Soguero-Ruiz, Filippo Maria Bianchi, Arthur Revhaug 외

The time series cluster kernel (TCK) provides a powerful tool for analysing multivariate time series subject to missing data. TCK is designed using an ensemble learning approach in which Bayesian mixture models form the …

Ensemble LearningImputationMissing ValuesTime Series+1

Deep Generative Pattern-Set Mixture Models for Nonignorable Missingness

2021-03-05 · Sahra Ghalebikesabi, Rob Cornish, Luke J. Kelly, Chris Holmes

We propose a variational autoencoder architecture to model both ignorable and nonignorable missing data using pattern-set mixtures as proposed by Little (1993). Our model explicitly learns to cluster the missing data int…

Imputation

Informative Label Missingness in Multiclass Classification Information Geometry and Excess Risk

2026-08-31 · Fariborz Setoudehtazang, Geoffrey J. McLachlan arxiv

Informative label missingness can change the usual efficiency ordering between completely and partially labelled classifiers because the pattern of missing labels may itself carry information about the classification mod…