paper-with-me

홈 › Papers

Dual Alignment Between Language Model Layers and Human Sentence Processing

2026-04-20 · Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki, Ethan Gotlieb Wilcox arxiv

A recent study (Kuribayashi et al., 2025) has shown that human sentence processing behavior, typically measured on syntactically unchallenging constructions, can be effectively modeled using surprisal from early layers of large language models (LLMs). This raises the question of whether such advantages of internal layers extend to more syntactically challenging constructions, where surprisal has been reported to underestimate human cognitive effort. In this paper, we begin by exploring internal layers that better estimate human cognitive effort observed in syntactic ambiguity processing in English. Our experiments show that, in contrast to naturalistic reading, later layers better estimate such a cognitive effort, but still underestimate the human data. This dual alignment sheds light on different modes of sentence processing in humans and LMs: naturalistic reading employs a somewhat weak prediction akin to earlier layers of LMs, while syntactically challenging processing requires more fully-contextualized representations, better modeled by later layers of LMs. Motivated by these findings, we also explore several probability-update measures using shallow and deep layers of LMs, showing a complementary advantage to single-layer's surprisal in reading time modeling.

📄 PDF Abstract BibTeX arXiv:2604.18563

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection

2025-09-19 · Maithili Joshi, Palash Nandi, Tanmoy Chakraborty arxiv

Large Language Models (LLMs) with safe-alignment training are powerful instruments with robust language comprehension capabilities. These models typically undergo meticulous alignment procedures involving human feedback …

Alignment is Localized: A Causal Probe into Preference Layers

2025-10-17 · Archie Chaudhury arxiv

Reinforcement Learning frameworks, particularly those utilizing human annotations, have become an increasingly popular method for preference fine-tuning, where the outputs of a language model are tuned to match a certain…

Reinforcement Learning

DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

2026-01-31 · Wenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han 외 arxiv

Few-shot learning (FSL) aims to generalize to novel categories with only a few samples. Recent approaches incorporate large language models (LLMs) to enrich visual representations with semantic embeddings derived from cl…

Reinforcement LearningFew-Shot Learning

Every word counts: A multilingual analysis of individual human alignment with model attention

2022-10-05 · Stephanie Brandl, Nora Hollenstein

Human fixation patterns have been shown to correlate strongly with Transformer-based attention. Those correlation analyses are usually carried out without taking into account individual differences between participants a…

Toward Modeling Player-Specific Chess Behaviors

2026-05-12 · Loris Sogliuzzo, Aloïs Rautureau, Eric Piette arxiv

While artificial intelligence has achieved superhuman performance in chess, developing models that accurately emulate the individualized decision-making styles of human players remains a significant challenge. Existing h…