paper-with-me

홈 › Papers

Machine Learning Capability: A standardized metric using case difficulty with applications to individualized deployment of supervised machine learning

2023-02-09 · Adrienne Kline, Joon Lee

Model evaluation is a critical component in supervised machine learning classification analyses. Traditional metrics do not currently incorporate case difficulty. This renders the classification results unbenchmarked for generalization. Item Response Theory (IRT) and Computer Adaptive Testing (CAT) with machine learning can benchmark datasets independent of the end-classification results. This provides high levels of case-level information regarding evaluation utility. To showcase, two datasets were used: 1) health-related and 2) physical science. For the health dataset a two-parameter IRT model, and for the physical science dataset a polytonomous IRT model, was used to analyze predictive features and place each case on a difficulty continuum. A CAT approach was used to ascertain the algorithms' performance and applicability to new data. This method provides an efficient way to benchmark data, using only a fraction of the dataset (less than 1%) and 22-60x more computationally efficient than traditional metrics. This novel metric, termed Machine Learning Capability (MLC) has additional benefits as it is unbiased to outcome classification and a standardized way to make model comparisons within and across datasets. MLC provides a metric on the limitation of supervised machine learning algorithms. In situations where the algorithm falls short, other input(s) are required for decision-making.

📄 PDF Abstract BibTeX arXiv:2302.04386

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationDecision Making

Similar Papers 제목 키워드 기반

ProteinNet: a standardized data set for machine learning of protein structure

2019-02-01 · Mohammed AlQuraishi

Rapid progress in deep learning has spurred its application to bioinformatics problems including protein structure prediction and design. In classic machine learning problems like computer vision, progress has been drive…

BIG-bench Machine LearningProtein Secondary Structure PredictionProtein Structure Prediction

RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers

2025-09-30 · Yifan Lu, Rixin Liu, Jiayi Yuan, Xingqi Cui 외 arxiv

Today's LLM ecosystem comprises a wide spectrum of models that differ in size, capability, and cost. No single model is optimal for all scenarios; hence, LLM routers have become essential for selecting the most appropria…

On the post-hoc Evaluation of PDE Discovery: A Multifaceted Challenge of Scientific Advancement

2026-07-26 · Baptiste Mathevon, Farah Cherfaoui, Amaury Habrard, Marc Sebban arxiv

Partial differential equation (PDE) discovery aims to identify from data the governing law of a physical system. Constituting a cornerstone of scientific advancement, it has become during the past decade a major line of …

Mind the Long Tail: Understanding the Difficulty of Delay Detection in Business Processes

2026-08-14 · Keyvan Amiri Elyasi, Lukas Kirchdorfer, Heiner Stuckenschmidt arxiv

The early detection of delayed cases in business processes is a critical capability for organizations. Predictive process monitoring (PPM) supports this task by using historical event logs to predict the remaining time o…

Multi-Dimensional Ability Diagnosis for Machine Learning Algorithms

2023-07-14 · Qi Liu, Zheng Gong, Zhenya Huang, Chuanren Liu 외

Machine learning algorithms have become ubiquitous in a number of applications (e.g. image classification). However, due to the insufficient measurement of traditional metrics (e.g. the coarse-grained Accuracy of each cl…

cognitive diagnosisDiagnosticimage-classificationImage Classification