paper-with-me

홈 › Papers

Probabilistic Predictions of People Perusing: Evaluating Metrics of Language Model Performance for Psycholinguistic Modeling

2020-09-08 · EMNLP (CMCL) 2020 11 · Yiding Hao, Simon Mendelsohn, Rachel Sterneck, Randi Martinez, Robert Frank

By positing a relationship between naturalistic reading times and information-theoretic surprisal, surprisal theory (Hale, 2001; Levy, 2008) provides a natural interface between language models and psycholinguistic models. This paper re-evaluates a claim due to Goodkind and Bicknell (2018) that a language model's ability to model reading times is a linear function of its perplexity. By extending Goodkind and Bicknell's analysis to modern neural architectures, we show that the proposed relation does not always hold for Long Short-Term Memory networks, Transformers, and pre-trained models. We introduce an alternate measure of language modeling performance called predictability norm correlation based on Cloze probabilities measured from human subjects. Our new metric yields a more robust relationship between language model quality and psycholinguistic modeling performance that allows for comparison between models with different training configurations.

📄 PDF Abstract BibTeX arXiv:2009.03954

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

XForecast: Evaluating Natural Language Explanations for Time Series Forecasting

2024-10-18 · Taha Aksu, Chenghao Liu, Amrita Saha, Sarah Tan 외

Time series forecasting aids decision-making, especially for stakeholders who rely on accurate predictions, making it very important to understand and explain these models to ensure informed decisions. Traditional explai…

Decision MakingTime SeriesTime Series Forecasting

Skimming, Locating, then Perusing: A Human-Like Framework for Natural Language Video Localization

2022-07-27 · Daizong Liu, Wei Hu

This paper addresses the problem of natural language video localization (NLVL). Almost all existing works follow the "only look once" framework that exploits a single model to directly capture the complex cross- and self…

From Classification Accuracy to Proper Scoring Rules: Elicitability of Probabilistic Top List Predictions

2023-01-27 · Johannes Resin

In the face of uncertainty, the need for probabilistic assessments has long been recognized in the literature on forecasting. In classification, however, comparative evaluation of classifiers often focuses on predictions…

Uncertainty Quantification

The Certainty Ratio $C_ρ$: a novel metric for assessing the reliability of classifier predictions

2024-11-04 · Jesus S. Aguilar-Ruiz

Evaluating the performance of classifiers is critical in machine learning, particularly in high-stakes applications where the reliability of predictions can significantly impact decision-making. Traditional performance m…

Decision Making

VisCUIT: Visual Auditor for Bias in CNN Image Classifier

2022-04-12 · CVPR 2022 1 · Seongmin Lee, Zijie J. Wang, Judy Hoffman, Duen Horng Chau

CNN image classifiers are widely used, thanks to their efficiency and accuracy. However, they can suffer from biases that impede their practical applications. Most existing bias investigation techniques are either inappl…

image-classificationImage Classification