paper-with-me

홈 › Papers

Recall, Robustness, and Lexicographic Evaluation

2023-02-22 · Fernando Diaz, Michael D. Ekstrand, Bhaskar Mitra

Although originally developed to evaluate sets of items, recall is often used to evaluate rankings of items, including those produced by recommender, retrieval, and other machine learning systems. The application of recall without a formal evaluative motivation has led to criticism of recall as a vague or inappropriate measure. In light of this debate, we reflect on the measurement of recall in rankings from a formal perspective. Our analysis is composed of three tenets: recall, robustness, and lexicographic evaluation. First, we formally define `recall-orientation' as the sensitivity of a metric to a user interested in finding every relevant item. Second, we analyze recall-orientation from the perspective of robustness with respect to possible content consumers and providers, connecting recall to recent conversations about fair ranking. Finally, we extend this conceptual and theoretical treatment of recall by developing a practical preference-based evaluation method based on lexicographic comparison. Through extensive empirical analysis across three recommendation tasks and 17 information retrieval tasks, we establish that our new evaluation method, lexirecall, has convergent validity (i.e., it is correlated with existing recall metrics) and exhibits substantially higher sensitivity in terms of discriminative power and stability in the presence of missing labels. Our conceptual, theoretical, and empirical analysis substantially deepens our understanding of recall and motivates its adoption through connections to robustness and fairness.

📄 PDF Abstract BibTeX arXiv:2302.11370

Code (1)

diazf/pref_eval 공식 구현

Tasks

FairnessInformation RetrievalMissing LabelsRetrievalSensitivity

Similar Papers 제목 키워드 기반

Best-Case Retrieval Evaluation: Improving the Sensitivity of Reciprocal Rank with Lexicographic Precision

2023-06-13 · Fernando Diaz

Across a variety of ranking tasks, researchers use reciprocal rank to measure the effectiveness for users interested in exactly one relevant item. Despite its widespread use, evidence suggests that reciprocal rank is bri…

RetrievalSensitivity

Finding a human-like classifier

2019-11-13 · Anonymous

There were many attempts to explain the trade-off between accuracy and adversarial robustness. However, there was no clear understanding of the behaviors of a robust classifier which has human-like robustne…

Adversarial RobustnessContinual Learning

Lexicographic Minimum-Violation Motion Planning using Signal Temporal Logic

2026-04-22 · Patrick Halder, Lothar Kiltz, Hannes Homburger, Johannes Reuter 외 arxiv

Motion planning for autonomous vehicles often requires satisfying multiple conditionally conflicting specifications. In situations where not all specifications can be met simultaneously, minimum-violation motion planning…

Autonomous VehiclesMotion Planning

Bounded Robustness in Reinforcement Learning via Lexicographic Objectives

2022-09-30 · Daniel Jarne Ornia, Licio Romao, Lewis Hammond, Manuel Mazo Jr. 외

Policy robustness in Reinforcement Learning may not be desirable at any cost: the alterations caused by robustness requirements from otherwise optimal policies should be explainable, quantifiable and formally verifiable.…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Training Word Sense Embeddings With Lexicon-based Regularization

2017-11-01 · IJCNLP 2017 11 · Luis Nieto-Pi{\~n}a, Richard Johansson

We propose to improve word sense embeddings by enriching an automatic corpus-based method with lexicographic data. Information from a lexicon is introduced into the learning algorithm{'}s objective function through a reg…

Word EmbeddingsWord Sense Disambiguation