paper-with-me

홈 › Papers

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English

2026-05-26 · Yuanhao Chen, Peter Chin arxiv

We show that LLMs encode syntactic distinctions not present in the Universal Dependencies (UD) tree distances that structural probes are trained to recover. On English wh-movement stimuli, we measure the probe distance between an embedded subject and its verb, whose UD tree distance is invariant across conditions. That distance is shorter than baseline when the embedded clause is finite and longer when it is infinitival -- a within-clause sign asymmetry present in all 13 models across four families we test, at a majority of layers. No account based on UD distance, linear order, or monotone structural complexity can produce a sign reversal, while Minimalist phase theory can. Holding the matrix verb fixed while varying only the complement type reproduces the same finite-infinitival ordering in every model, ruling out a lexical-semantic explanation. The cross-clause pair separately reproduces the phase-count ordering of earlier work, validating the probes. A single activation patch, interchanging the clause-selecting matrix verb, moves the two pairs in opposite directions in 8 of 13 models. Because the causal intervention leaves the target string unchanged, it also rules out the clause-length and surface-cue explanations. Together these results reveal syntactic structure in LLMs beyond the UD probe target, and a Minimalist phase account is consistent with it.

📄 PDF Abstract BibTeX arXiv:2605.26431

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Universal and Independent: Multilingual Probing Framework for Exhaustive Model Interpretation and Evaluation

2022-10-24 · Oleg Serikov, Vitaly Protasov, Ekaterina Voloshina, Viktoria Knyazkova 외

Linguistic analysis of language models is one of the ways to explain and describe their reasoning, weaknesses, and limitations. In the probing part of the model interpretability research, studies concern individual langu…

Probing Language Models

Do Neural Language Models Show Preferences for Syntactic Formalisms?

2020-04-29 · ACL 2020 6 · Artur Kulmizev, Vinit Ravishankar, Mostafa Abdou, Joakim Nivre

Recent work on the interpretability of deep neural language models has concluded that many properties of natural language syntax are encoded in their representational spaces. However, such studies often suffer from limit…

Probing sentence embeddings for structure-dependent tense

2018-11-01 · WS 2018 11 · Geoff Bacon, Terry Regier

Learning universal sentence representations which accurately model sentential semantic content is a current goal of natural language processing research. A prominent and successful approach is to train recurrent neural n…

SentenceSentence Embeddings

Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian

2026-03-26 · Giuseppe Samo, Paola Merlo arxiv

This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Italian. We use Blackbird Language Matrices (…

What can Neural Referential Form Selectors Learn?

2021-08-15 · INLG (ACL) 2021 8 · Guanyi Chen, Fahime Same, Kees Van Deemter

Despite achieving encouraging results, neural Referring Expression Generation models are often thought to lack transparency. We probed neural Referential Form Selection (RFS) models to find out to what extent the linguis…

FormPositionReferring ExpressionReferring expression generation+1