paper-with-me

홈 › Papers

Can LLMs capture stable human-generated sentence entropy measures?

2026-02-04 · Estrella Pivel-Villanueva, Elisabeth Frederike Sterner, Franziska Knolle arxiv

Predicting upcoming words is a core mechanism of language comprehension and may be quantified using Shannon entropy. There is currently no empirical consensus on how many human responses are required to obtain stable and unbiased entropy estimates at the word level. Moreover, large language models (LLMs) are increasingly used as substitutes for human norming data, yet their ability to reproduce stable human entropy remains unclear. Here, we address both issues using two large publicly available cloze datasets in German 1 and English 2. We implemented a bootstrap-based convergence analysis that tracks how entropy estimates stabilize as a function of sample size. Across both languages, more than 97% of sentences reached stable entropy estimates within the available sample sizes. 90% of sentences converged after 111 responses in German and 81 responses in English, while low-entropy sentences (<1) required as few as 20 responses and high-entropy sentences (>2.5) substantially more. These findings provide the first direct empirical validation for common norming practices and demonstrate that convergence critically depends on sentence predictability. We then compared stable human entropy values with entropy estimates derived from several LLMs, including GPT-4o, using both logit-based probability extraction and sampling-based frequency estimation, GPT2-xl/german-GPT-2, RoBERTa Base/GottBERT, and LLaMA 2 7B Chat. GPT-4o showed the highest correspondence with human data, although alignment depended strongly on the extraction method and prompt design. Logit-based estimates minimized absolute error, whereas sampling-based estimates were better in capturing the dispersion of human variability. Together, our results establish practical guidelines for human norming and show that while LLMs can approximate human entropy, they are not interchangeable with stable human-derived distributions.

📄 PDF Abstract BibTeX arXiv:2602.04570

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language models align with human judgments on key grammatical constructions

2024-01-19 · Jennifer Hu, Kyle Mahowald, Gary Lupyan, Anna Ivanova 외

Do large language models (LLMs) make human-like linguistic generalizations? Dentella et al. (2023) ("DGL") prompt several LLMs ("Is the following sentence grammatically correct in English?") to elicit grammaticality judg…

Sentence

LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning

2024-02-02 · Rongsheng Wang, Haoming Chen, Ruizhe Zhou, Han Ma 외

ChatGPT and other general large language models (LLMs) have achieved remarkable success, but they have also raised concerns about the misuse of AI-generated texts. Existing AI-generated text detection models, such as bas…

SentenceText Detection

SeqXGPT: Sentence-Level AI-Generated Text Detection

2023-10-13 · Pengyu Wang, Linyang Li, Ke Ren, Botian Jiang 외

Widely applied large language models (LLMs) can generate human-like content, raising concerns about the abuse of LLMs. Therefore, it is important to build strong AI-generated text (AIGT) detectors. Current works only con…

SentenceText Detection

An Exploratory Study on LLM-Generated Code and Comments in Code Repositories

2026-07-02 · Yongyi Ji, Jiaji Wang, Yi Zhou, Fuxiang Chen 외 arxiv

The use of LLMs in software development has become increasingly widespread on tasks such as code generation and summarization. Reports from large technology companies showed that around 20% to 30% of their code are gener…

Code Generation

Do Large Language Models know who did what to whom?

2025-04-23 · Joseph M. Denning, Xiaohan, Guo, Bryor Snefjella 외

Large Language Models (LLMs) are commonly criticized for not understanding language. However, many critiques focus on cognitive abilities that, in humans, are distinct from language processing. Here, we instead study a k…

Sentence