paper-with-me

홈 › Papers

Psychometric Predictive Power of Large Language Models

2023-11-13 · Tatsuki Kuribayashi, Yohei Oseki, Timothy Baldwin

Instruction tuning aligns the response of large language models (LLMs) with human preferences. Despite such efforts in human--LLM alignment, we find that instruction tuning does not always make LLMs human-like from a cognitive modeling perspective. More specifically, next-word probabilities estimated by instruction-tuned LLMs are often worse at simulating human reading behavior than those estimated by base LLMs. In addition, we explore prompting methodologies for simulating human reading behavior with LLMs. Our results show that prompts reflecting a particular linguistic hypothesis improve psychometric predictive power, but are still inferior to small base models. These findings highlight that recent advancements in LLMs, i.e., instruction tuning and prompting, do not offer better estimates than direct probability measurements from base LLMs in cognitive modeling. In other words, pure next-word probability remains a strong predictor for human reading behavior, even in the age of LLMs.

📄 PDF Abstract BibTeX arXiv:2311.07484

Code (1)

kuribayashi4/llm-cognitive-modeling 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

On the Predictive Power of Neural Language Models for Human Real-Time Comprehension Behavior

2020-06-02 · Ethan Gotlieb Wilcox, Jon Gauthier, Jennifer Hu, Peng Qian 외

Human reading behavior is tuned to the statistics of natural language: the time it takes human subjects to read a word can be predicted from estimates of the word's probability in context. However, it remains an open que…

Open-Ended Question Answering

Reverse-Engineering the Reader

2024-10-16 · Samuel Kiegeland, Ethan Gotlieb Wilcox, Afra Amini, David Robert Reich 외

Numerous previous studies have sought to determine to what extent language models, pretrained on natural language text, can serve as useful models of human cognition. In this paper, we are interested in the opposite ques…

Language ModelingLanguage Modelling

Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models

2024-09-18 · Haoran Ye, Yuhang Xie, Yuanyi Ren, Hanjun Fang 외

Human values and their measurement are long-standing interdisciplinary inquiry. Recent advances in AI have sparked renewed interest in this area, with large language models (LLMs) emerging as both tools and subjects of v…

AI Psychometrics: Evaluating the Psychological Reasoning of Large Language Models with Psychometric Validities

2026-03-11 · Yibai Li, Xiaolin Lin, Zhenghui Sha, Zhiye Jin 외 arxiv

The immense number of parameters and deep neural networks make large language models (LLMs) rival the complexity of human brains, which also makes them opaque ``black box'' systems that are challenging to evaluate and in…

Language models emulate certain cognitive profiles: An investigation of how predictability measures interact with individual differences

2024-06-07 · Patrick Haller, Lena S. Bolliger, Lena A. Jäger

To date, most investigations on surprisal and entropy effects in reading have been conducted on the group level, disregarding individual differences. In this work, we revisit the predictive power of surprisal and entropy…