Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation
Commonly used evaluation measures including Recall, Precision, F-Measure and Rand Accuracy are biased and should not be used without clear understanding of the biases, and corresponding identification of chance or base case levels of the statistic. Using these measures a system that performs worse in the objective sense of Informedness, can appear to perform better under any of these commonly used measures. We discuss several concepts and measures that reflect the probability that prediction is informed versus chance. Informedness and introduce Markedness as a dual measure for the probability that prediction is marked versus chance. Finally we demonstrate elegant connections between the concepts of Informedness, Markedness, Correlation and Significance as well as their intuitive relationships with Recall and Precision, and outline the extension from the dichotomous case to the general multi-class case.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Evaluation Evaluation a Monte Carlo study
Over the last decade there has been increasing concern about the biases embodied in traditional evaluation methods for Natural Language Processing/Learning, particularly methods borrowed from Information Retrieval. Witho…
Information RetrievalRetrievalWe Need to Talk About Classification Evaluation Metrics in NLP
In Natural Language Processing (NLP) classification tasks such as topic categorisation and sentiment analysis, model generalizability is generally measured with standard metrics such as Accuracy, F-Measure, or AUC-ROC. T…
DiversityMachine TranslationNatural Language UnderstandingQuestion Answering+1Modeling Markedness with a Split-and-Merger Model of Sound Change
The concept of {`}markedness{'} has been influential in phonology for almost a century. Theoretical phonology has found it useful to describe some segments as more {`}marked{'} than others, referring to a cluster of lang…
Measures and Meta-Measures for the Supervised Evaluation of Image Segmentation
This paper tackles the supervised evaluation of image segmentation algorithms. First, it surveys and structures the measures used to compare the segmentation results with a ground truth database; and proposes a new measu…
Image SegmentationSegmentationSemantic SegmentationStatistical Evaluation of Anomaly Detectors for Sequences
Although precision and recall are standard performance measures for anomaly detection, their statistical properties in sequential detection settings are poorly understood. In this work, we formalize a notion of precision…
Anomaly Detection