The curious case of the test set AUROC
Whilst the size and complexity of ML models have rapidly and significantly increased over the past decade, the methods for assessing their performance have not kept pace. In particular, among the many potential performance metrics, the ML community stubbornly continues to use (a) the area under the receiver operating characteristic curve (AUROC) for a validation and test cohort (distinct from training data) or (b) the sensitivity and specificity for the test data at an optimal threshold determined from the validation ROC. However, we argue that considering scores derived from the test ROC curve alone gives only a narrow insight into how a model performs and its ability to generalise.
Code (1)
Tasks
SensitivitySpecificitySimilar Papers 제목 키워드 기반
The Curious Case of Metonymic Verbs: A Distributional Characterization
Word Embeddings vs Word Types for Sequence Labeling: the Curious Case of CV Parsing
Evaluation of MRI to ultrasound registration methods for brain shift correction: The CuRIOUS2018 Challenge
In brain tumor surgery, the quality and safety of the procedure can be impacted by intra-operative tissue deformation, called brain shift. Brain shift can move the surgical targets and other vital structures such as bloo…
Image RegistrationCurious Explorer: a provable exploration strategy in Policy Learning
Having access to an exploring restart distribution (the so-called wide coverage assumption) is critical with policy gradient methods. This is due to the fact that, while the objective function is insensitive to updates i…
Policy Gradient MethodsArtificial Intelligence-Based Triaging of Cutaneous Melanocytic Lesions
Pathologists are facing an increasing workload due to a growing volume of cases and the need for more comprehensive diagnoses. Aiming to facilitate workload reduction and faster turnaround times, we developed an artifici…
whole slide images