paper-with-me

홈 › Papers

Structural Persistence in Language Models: Priming as a Window into Abstract Language Representations

2021-09-30 · Arabella Sinclair, Jaap Jumelet, Willem Zuidema, Raquel Fernández

We investigate the extent to which modern, neural language models are susceptible to structural priming, the phenomenon whereby the structure of a sentence makes the same structure more probable in a follow-up sentence. We explore how priming can be used to study the potential of these models to learn abstract structural information, which is a prerequisite for good performance on tasks that require natural language understanding skills. We introduce a novel metric and release Prime-LM, a large corpus where we control for various linguistic factors which interact with priming strength. We find that Transformer models indeed show evidence of structural priming, but also that the generalisations they learned are to some extent modulated by semantic information. Our experiments also show that the representations acquired by the models may not only encode abstract sequential structure but involve certain level of hierarchical syntactic information. More generally, our study shows that the priming paradigm is a useful, additional tool for gaining insights into the capacities of language models and opens the door to future priming-based investigations that probe the model's internal states.

📄 PDF Abstract BibTeX arXiv:2109.14989

Code (1)

dmg-illc/prime-lm 공식 구현 pytorch

Tasks

Natural Language UnderstandingSentence

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

Towards Human Cognition: Visual Context Guides Syntactic Priming in Fusion-Encoded Models

2025-02-24 · Bushi Xiao, Michael Bennie, Jayetri Bardhan, Daisy Zhe Wang

We introduced PRISMATIC, the first multimodal structural priming dataset, and proposed a reference-free evaluation metric that assesses priming effects without predefined target sentences. Using this metric, we construct…

Semantic-Oriented Unlabeled Priming for Large-Scale Language Models

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Due to the high costs associated with finetuning large language models, various recent works propose to adapt them to specific tasks without any parameter updates through in-context learning. Unfortunately, for in-contex…

In-Context Learning

Semantic-Oriented Unlabeled Priming for Large-Scale Language Models

2022-02-12 · Yanchen Liu, Timo Schick, Hinrich Schütze

Due to the high costs associated with finetuning large language models, various recent works propose to adapt them to specific tasks without any parameter updates through in-context learning. Unfortunately, for in-contex…

In-Context Learning

Do Language Models Exhibit Human-like Structural Priming Effects?

2024-06-07 · Jaap Jumelet, Willem Zuidema, Arabella Sinclair

We explore which linguistic factors -- at the sentence and token level -- play an important role in influencing language model predictions, and investigate whether these are reflective of results found in humans and huma…

Language ModelingLanguage ModellingSentence

Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models

2023-11-15 · James A. Michaelov, Catherine Arnett, Tyler A. Chang, Benjamin K. Bergen

Abstract grammatical knowledge - of parts of speech and grammatical patterns - is key to the capacity for linguistic generalization in humans. But how abstract is grammatical knowledge in large language models? In the hu…

Sentence