paper-with-me

홈 › Papers

Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

2021-08-01 · ACL 2021 5 · Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, Jordan Boyd-Graber

Leaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). Rather than replace leaderboards, we advocate a re-imagining so that they better highlight if and where progress is made. Building on educational testing, we create a Bayesian leaderboard model where latent subject skill and latent item difficulty predict correct responses. Using this model, we analyze the ranking reliability of leaderboards. Afterwards, we show the model can guide what to annotate, identify annotation errors, detect overfitting, and identify informative examples. We conclude with recommendations for future benchmark tasks.

📄 PDF Abstract BibTeX

Code (1)

jplalor/py-irt 공식 구현 pytorch

Similar Papers 제목 키워드 기반

On Quantitative Evaluations of Counterfactuals

2021-10-30 · Frederik Hvilshøj, Alexandros Iosifidis, Ira Assent

As counterfactual examples become increasingly popular for explaining decisions of deep learning models, it is essential to understand what properties quantitative evaluation metrics do capture and equally important what…

counterfactual

Adversarial Over-Sensitivity and Over-Stability Strategies for Dialogue Models

2018-09-06 · CONLL 2018 10 · Tong Niu, Mohit Bansal

We present two categories of model-agnostic adversarial strategies that reveal the weaknesses of several generative, task-oriented dialogue models: Should-Not-Change strategies that evaluate over-sensitivity to small and…

Sensitivity

Para-active learning

2013-10-30 · Alekh Agarwal, Leon Bottou, Miroslav Dudik, John Langford

Training examples are not all equally informative. Active learning strategies leverage this observation in order to massively reduce the number of examples that need to be labeled. We leverage the same observation to bui…

Active Learning

AI Testing Should Account for Sophisticated Strategic Behaviour

2025-08-19 · Vojtech Kovarik, Eric Olav Chen, Sami Petersen, Alexis Ghersengorin 외 arxiv

This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility that AI systems understand their circum…

Evaluating Paraphrastic Robustness in Textual Entailment Models

2023-06-29 · Dhruv Verma, Yash Kumar Lal, Shreyashee Sinha, Benjamin Van Durme 외

We present PaRTE, a collection of 1,126 pairs of Recognizing Textual Entailment (RTE) examples to evaluate whether models are robust to paraphrasing. We posit that if RTE models understand language, their predictions sho…

Natural Language InferenceRTE