paper-with-me

Papers

Stratified Sampling for Extreme Multi-Label Data

2021-03-05 · Maximillian Merrillees, Lan Du

Extreme multi-label classification (XML) is becoming increasingly relevant in the era of big data. Yet, there is no method for effectively generating stratified partitions of XML datasets. Instead, researchers typically rely on provided test-train splits that, 1) aren't always representative of the entire dataset, and 2) are missing many of the labels. This can lead to poor generalization ability and unreliable performance estimates, as has been established in the binary and multi-class settings. As such, this paper presents a new and simple algorithm that can efficiently generate stratified partitions of XML datasets with millions of unique labels. We also examine the label distributions of prevailing benchmark splits, and investigate the issues that arise from using unrepresentative subsets of data for model development. The results highlight the difficulty of stratifying XML data, and demonstrate the importance of using stratified partitions for training and evaluation.

📄 PDF Abstract BibTeX arXiv:2103.03494

Code (2)

maxitron93/stratified_sampling_for_XML 공식 구현
arbasher/straSplit

Tasks

Extreme Multi-Label ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

Similar Papers 제목 키워드 기반

Label-Efficient Monitoring of Classification Models via Stratified Importance Sampling

2026-01-29 · Lupo Marsigli, Angel Lopez de Haro arxiv

Monitoring the performance of classification models in production is critical yet challenging due to strict labeling budgets, one-shot batch acquisition of labels and extremely low error rates. We propose a general frame…

Near Optimal Stratified Sampling

2019-06-26 · Tiancheng Yu, Xiyu Zhai, Suvrit Sra

The performance of a machine learning system is usually evaluated by using i.i.d.\ observations with true labels. However, acquiring ground truth labels is expensive, while obtaining unlabeled samples may be cheaper. Str…

Cancer Prognosis Prediction Using Balanced Stratified Sampling

2014-03-12 · J S Saleema, N Bhagawathi, S Monica, P Deepa Shenoy 외

High accuracy in cancer prediction is important to improve the quality of the treatment and to improve the rate of survivability of patients. As the data volume is increasing rapidly in the healthcare research, the analy…

General ClassificationPredictionPrognosis

Ensemble Learning with Sparse Hypercolumns

2026-03-06 · Julia Dietlmeier, Vayangi Ganepola, Oluwabukola G. Adegboro, Mayug Maniparambil 외 arxiv

Directly inspired by findings in biological vision, high-dimensional hypercolumns are feature vectors built by concatenating multi-scale activations of convolutional neural networks for a single image pixel location. Tog…

Image SegmentationEnsemble Learning

Stratified Prediction-Powered Inference for Hybrid Language Model Evaluation

2024-06-06 · Adam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra 외

Prediction-powered inference (PPI) is a method that improves statistical estimates based on limited human-labeled data. PPI achieves this by combining small amounts of human-labeled data with larger amounts of data label…

Language Model EvaluationLanguage ModelingLanguage Modellingvalid