paper-with-me

홈 › Papers

Predict the Next Word: Humans exhibit uncertainty in this task and language models _____

2024-02-27 · Evgenia Ilia, Wilker Aziz

Language models (LMs) are statistical models trained to assign probability to human-generated text. As such, it is reasonable to question whether they approximate linguistic variability exhibited by humans well. This form of statistical assessment is difficult to perform at the passage level, for it requires acceptability judgements (i.e., human evaluation) or a robust automated proxy (which is non-trivial). At the word level, however, given some context, samples from an LM can be assessed via exact matching against a prerecorded dataset of alternative single-word continuations of the available context. We exploit this fact and evaluate the LM's ability to reproduce variability that humans (in particular, a population of English speakers) exhibit in the 'next word prediction' task. This can be seen as assessing a form of calibration, which, in the context of text classification, Baan et al. (2022) termed calibration to human uncertainty. We assess GPT2, BLOOM and ChatGPT and find that they exhibit fairly low calibration to human uncertainty. We also verify the failure of expected calibration error (ECE) to reflect this, and as such, advise the community against relying on it in this setting.

📄 PDF Abstract BibTeX arXiv:2402.17527

Code (2)

evgeniael/predict_next_word 공식 구현 pytorch
evgeniael/probar pytorch

Tasks

text-classificationText Classification

Methods 이 논문이 사용한 방법론

BLOOM BLOOM is a decoder-only Transformer language model that was trained on the ROOTS corpus, a dataset comprising hundreds of sources in 46 natural and 13 programming languages…

Similar Papers 제목 키워드 기반

Humans and language models diverge when predicting repeating text

2023-10-10 · Aditya R. Vaidya, Javier Turek, Alexander G. Huth

Language models that are trained on the next-word prediction task have been shown to accurately model human behavior in word prediction and reading speed. In contrast with these findings, we present a scenario in which t…

In-Context Learning

The human unlikeness of neural language models in next-word prediction

2020-07-01 · WS 2020 7 · Cass Jacobs, ra L., Arya D. McCarthy

The training objective of unidirectional language models (LMs) is similar to a psycholinguistic benchmark known as the cloze task, which measures next-word predictability. However, LMs lack the rich set of experiences th…

To model human linguistic prediction, make LLMs less superhuman

2025-10-01 · Byung-Doh Oh, Tal Linzen arxiv

When we read, we make predictions about upcoming words; these predictions influence our reading behavior. The success of large language models (LLMs), which, like humans, make predictions about upcoming words, has motiva…

Explaining Predictive Uncertainty by Looking Back at Model Explanations

2022-01-11 · Hanjie Chen, Wanyu Du, Yangfeng Ji

Predictive uncertainty estimation of pre-trained language models is an important measure of how likely people can trust their predictions. However, little is known about what makes a model prediction uncertain. Explainin…

Decision MakingNatural Language InferenceParaphrase IdentificationPrediction+1

Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis

2026-05-14 · William Timkey, Brian Dillon, Tal Linzen arxiv

Surprisal theory posits that the processing difficulty of a word is determined by its predictability in context, offering a potential link between human sentence processing and next-word predictions from language models.…