paper-with-me

홈 › Papers

Ensemble- and Distance-Based Feature Ranking for Unsupervised Learning

2020-11-23 · Matej Petković, Dragi Kocev, Blaž Škrlj, Sašo Džeroski

In this work, we propose two novel (groups of) methods for unsupervised feature ranking and selection. The first group includes feature ranking scores (Genie3 score, RandomForest score) that are computed from ensembles of predictive clustering trees. The second method is URelief, the unsupervised extension of the Relief family of feature ranking algorithms. Using 26 benchmark data sets and 5 baselines, we show that both the Genie3 score (computed from the ensemble of extra trees) and the URelief method outperform the existing methods and that Genie3 performs best overall, in terms of predictive power of the top-ranked features. Additionally, we analyze the influence of the hyper-parameters of the proposed methods on their performance, and show that for the Genie3 score the highest quality is achieved by the most efficient parameter configuration. Finally, we propose a way of discovering the location of the features in the ranking, which are the most relevant in reality.

📄 PDF Abstract BibTeX arXiv:2011.11679

Code (1)

Petkomat/unsupervised_ranking 공식 구현

Tasks

Clustering

Similar Papers 제목 키워드 기반

Cross Device Matching for Online Advertising with Neural Feature Ensembles : First Place Solution at CIKM Cup 2016

2016-10-23 · Minh C. Phan, Yi Tay, Tuan-Anh Nguyen Pham

We describe the 1st place winning approach for the CIKM Cup 2016 Challenge. In this paper, we provide an approach to reasonably identify same users across multiple devices based on browsing logs. Our approach regards a c…

ClassificationGeneral Classification

Unsupervised Ensemble Ranking of Terms in Electronic Health Record Notes Based on Their Importance to Patients

2017-03-01 · Jinying Chen, Hong Yu

Background: Electronic health record (EHR) notes contain abundant medical jargon that can be difficult for patients to comprehend. One way to help patients is to reduce information overload and help them focus on medical…

Re-ranking Person Re-identification with k-reciprocal Encoding

2017-01-29 · CVPR 2017 7 · Zhun Zhong, Liang Zheng, Donglin Cao, Shaozi Li

When considering person re-identification (re-ID) as a retrieval process, re-ranking is a critical step to improve its accuracy. Yet in the re-ID community, limited effort has been devoted to re-ranking, especially those…

Person Re-IdentificationRe-RankingRetrieval

Distractor Generation for Multiple Choice Questions Using Learning to Rank

2018-06-01 · WS 2018 6 · Chen Liang, Xiao Yang, Neisarg Dave, Drew Wham 외

We investigate how machine learning models, specifically ranking models, can be used to select useful distractors for multiple choice questions. Our proposed models can learn to select distractors that resemble those in …

BIG-bench Machine LearningDistractor GenerationEnsemble LearningLearning-To-Rank+1

NLP-CIC @ DIACR-Ita: POS and Neighbor Based Distributional Models for Lexical Semantic Change in Diachronic Italian Corpora

2020-11-07 · Jason Angel, Carlos A. Rodriguez-Diaz, Alexander Gelbukh, Sergio Jimenez

We present our systems and findings on unsupervised lexical semantic change for the Italian language in the DIACR-Ita shared-task at EVALITA 2020. The task is to determine whether a target word has evolved its meaning wi…

POS