What the F-measure doesn't measure: Features, Flaws, Fallacies and Fixes
The F-measure or F-score is one of the most commonly used single number measures in Information Retrieval, Natural Language Processing and Machine Learning, but it is based on a mistake, and the flawed assumptions render it unsuitable for use in most contexts! Fortunately, there are better alternatives.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningInformation RetrievalRetrievalSimilar Papers 제목 키워드 기반
Measuring Calibration in Deep Learning
Overconfidence and underconfidence in machine learning classifiers is measured by calibration: the degree to which the probabilities predicted for each class match the accuracy of the classifier on that prediction. How o…
Deep LearningSarcasm SIGN: Interpreting Sarcasm with Sentiment Based Monolingual Machine Translation
Sarcasm is a form of speech in which speakers say the opposite of what they truly mean in order to convey a strong sentiment. In other words, "Sarcasm is the giant chasm between what I say, and the person who doesn't get…
Machine TranslationTranslationA Deep Learning Approach To Estimation Using Measurements Received Over a Network
We propose a novel deep neural network (DNN) based approximation architecture to learn estimates of measurements. We detail an algorithm that enables training of the DNN. The DNN estimator only uses measurements, if and …
Choose Your Lenses: Flaws in Gender Bias Evaluation
Considerable efforts to measure and mitigate gender bias in recent years have led to the introduction of an abundance of tasks, datasets, and metrics used in this vein. In this position paper, we assess the current parad…
"What if she doesn't feel the same?" What Happens When We Ask AI for Relationship Advice
Large Language Models (LLMs) are increasingly being used to provide support and advice in personal domains such as romantic relationships, yet little is known about user perceptions of this type of advice. This study inv…