The mRMR variable selection method: a comparative study for functional data
The use of variable selection methods is particularly appealing in
statistical problems with functional data. The obvious general criterion for
variable selection is to choose the most representative' or most relevant'
variables. However, it is also clear that a purely relevance-oriented criterion
could lead to select many redundant variables. The mRMR (minimum Redundance
Maximum Relevance) procedure, proposed by Ding and Peng (2005) and Peng et al.
(2005) is an algorithm to systematically perform variable selection, achieving
a reasonable trade-off between relevance and redundancy. In its original form,
this procedure is based on the use of the so-called mutual information
criterion to assess relevance and redundancy. Keeping the focus on functional
data problems, we propose here a modified version of the mRMR method, obtained
by replacing the mutual information by the new association measure (called
distance correlation) suggested by Sz\'ekely et al. (2007). We have also
performed an extensive simulation study, including 1600 functional experiments
(100 functional models $\times$ 4 sample sizes $\times$ 4 classifiers) and
three real-data examples aimed at comparing the different versions of the mRMR
methodology. The results are quite conclusive in favor of the new proposed
alternative.
Code (0)
등록된 구현이 없습니다.
Tasks
Variable SelectionSimilar Papers 제목 키워드 기반
varrank: an R package for variable ranking based on mutual information with applications to observed systemic datasets
This article describes the R package varrank. It has a flexible implementation of heuristic approaches which perform variable ranking based on mutual information. The package is particularly suitable for exploring multiv…
Dimensionality ReductionFeature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy
Feature selection is an important problem for pattern classification systems. We study how to select good features according to the maximal statistical dependency criterion based on mutual information. Because of the dif…
Classificationfeature selectionKGroups: A Versatile Univariate Max-Relevance Min-Redundancy Feature Selection Algorithm for High-dimensional Biological Data
This paper proposes a new univariate filter feature selection (FFS) algorithm called KGroups. The majority of work in the literature focuses on investigating the relevance or redundancy estimations of feature selection (…
BoMGene: Integrating Boruta-mRMR feature selection for enhanced Gene expression classification
Feature selection is a crucial step in analyzing gene expression data, enhancing classification performance, and reducing computational costs for high-dimensional datasets. This paper proposes BoMGene, a hybrid feature s…
Maximum Relevance and Minimum Redundancy Feature Selection Methods for a Marketing Machine Learning Platform
In machine learning applications for online product offerings and marketing strategies, there are often hundreds or thousands of features available to build such models. Feature selection is one essential method in such …
BIG-bench Machine Learningfeature selectionMarketing