paper-with-me

홈 › Papers

When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection

2026-06-30 · Jesus S. Aguilar-Ruiz arxiv

Feature rankings are widely used in supervised feature selection because they are simple, scalable and easy to interpret. Variables are first ranked by a relevance score, and a subset is then obtained by retaining the top-ranked variables. Although the first stage has been extensively studied, the second is often governed by an arbitrary cardinality, an empirical threshold or cross-validation, without a direct interpretation. This raises a basic question: given a feature ranking, when is there enough accumulated class-separation evidence to stop selecting features? This paper develops a distributional framework for transforming supervised feature rankings into class-independent subsets through an explicit risk-calibrated stopping rule. For each variable and each pair of classes, marginal separation is measured by the Bhattacharyya coefficient between the corresponding class-conditional distributions. The proposed method selects a single global subset shared by all classes by retaining the shortest prefix of a ranking whose residual product overlap falls below a prescribed threshold for every relevant class contrast. We derive binary and multiclass Bayes-risk bounds for the labelled product marginal problem, and obtain prior-dependent and prior-free calibrations of the residual-overlap threshold from a target all-pairs risk level. An empirical comparison on high-dimensional genomic datasets illustrates that the rule can reduce tens of thousands of variables to a few dozen while maintaining predictive performance statistically comparable to the all-features baseline. As the stopping rule only requires one-dimensional marginal overlap estimates and scans a precomputed ranking, it is well suited to very high-dimensional settings where exhaustive subset search is infeasible and interpretable truncation of feature rankings is essential.

📄 PDF Abstract BibTeX arXiv:2606.31686

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Determining Winners in Elections with Absent Votes

2023-10-11 · Qishen Han, Amélie Marian, Lirong Xia

An important question in elections is the determine whether a candidate can be a winner when some votes are absent. We study this determining winner with the absent votes (WAV) problem when the votes are top-truncated. W…

scoring rule

Beyond Logit Adjustment: A Residual Decomposition Framework for Long-Tailed Reranking

2026-04-02 · Zhanliang Wang, Hongzhuo Chen, Quan Minh Nguyen, Mian Umair Ahsan 외 arxiv

Long-tailed classification, where a small number of frequent classes dominate many rare ones, remains challenging because models systematically favor frequent classes at inference time. Existing post-hoc methods such as …

Image ClassificationScene Recognition

New fairness criteria for truncated ballots in multi-winner ranked-choice elections

2024-08-07 · Adam Graham-Squire, Matthew I. Jones, David McCune

In real-world elections where voters cast preference ballots, voters often provide only a partial ranking of the candidates. Despite this empirical reality, prior social choice literature frequently analyzes fairness cri…

Fairness

Ranking with Ties based on Noisy Performance Data

2024-05-28 · Aravind Sankaran, Lars Karlsson, Paolo Bientinesi

We consider the problem of ranking a set of objects based on their performance when the measurement of said performance is subject to noise. In this scenario, the performance is measured repeatedly, resulting in a range …

Efficient Parameter Estimation of Truncated Boolean Product Distributions

2020-07-05 · Dimitris Fotakis, Alkis Kalavasis, Christos Tzamos

We study the problem of estimating the parameters of a Boolean product distribution in $d$ dimensions, when the samples are truncated by a set $S \subset \{0, 1\}^d$ accessible through a membership oracle. This is the fi…

parameter estimation