paper-with-me

홈 › Papers

Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions

2026-05-29 · Wesley Scivetti, Ethan Wilcox, Nathan Schneider, Kanishka Misra, Leonie Weissweiler arxiv

Grasping the semantics of rare constructions (form-meaning pairings) has been shown to be a challenging problem that has currently only been solved by the largest LLMs. It remains an open question if open-source models have robust constructional understanding, and if so, what learning dynamics underlie the acquisition of this knowledge. Focusing on a set of rare Paired-Focus constructions in English (e.g. "let alone", "much less"), we construct a novel dataset to test their meanings using both scalar adjectival semantics and general world knowledge. Testing a wide range of models differing in parameter count, architecture, and pretraining dataset size, we find that several modestly sized models are sensitive to both the forms and the meanings of Paired-Focus constructions, though models trained on human-scale data fail at all meaning evaluations. Turning to training dynamics for a set of open-checkpoint models, we find that Paired-Focus understanding emerges later in training than Paired-Focus syntactic knowledge, and that learning of Paired-Focus semantics is correlated with gains in some domains of world knowledge. Overall, our empirical results support the conclusion that modestly sized open-source models can grasp the rare Paired-Focus constructions, and demonstrate a connection between knowledge of Paired-Focus constructions and other meaning domains.

📄 PDF Abstract BibTeX arXiv:2605.31586

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The D\'emonette Lexical Database: between Constructional Semantics and Word Formation (La base lexicale D\'emonette : entre s\'emantique constructionnelle et morphologie d\'erivationnelle) [in French]

2014-07-01 · JEPTALNRECITAL 2014 7 · Nabil Hathout, Fiammetta Namer

CxGBERT: BERT meets Construction Grammar

2020-11-09 · COLING 2020 8 · Harish Tayyar Madabushi, Laurence Romain, Dagmar Divjak, Petar Milin

While lexico-semantic elements no doubt capture a large amount of linguistic information, it has been argued that they do not capture all information contained in text. This assumption is central to constructionist appro…

Incorporating Syntax and Frame Semantics in Neural Network for Machine Reading Comprehension

2020-12-01 · COLING 2020 8 · Shaoru Guo, Yong Guan, Ru Li, XiaoLi Li 외

Machine reading comprehension (MRC) is one of the most critical yet challenging tasks in natural language understanding(NLU), where both syntax and semantics information of text are essential components for text understa…

Machine Reading ComprehensionNatural Language UnderstandingReading Comprehension

Disentangling Semantics and Syntax in Sentence Embeddings with Pre-trained Language Models

2021-04-11 · NAACL 2021 4 · James Y. Huang, Kuan-Hao Huang, Kai-Wei Chang

Pre-trained language models have achieved huge success on a wide range of NLP tasks. However, contextual representations from pre-trained models contain entangled semantic and syntactic information, and therefore cannot …

Semantic SimilaritySemantic Textual SimilaritySentenceSentence Embedding+2

Dependency resolution and semantic mining using Tree Adjoining Grammars for Tamil Language

2017-04-19 · Vijay Krishna Menon, S. Rajendran, M Anandkumar, K. P. Soman

Tree adjoining grammars (TAGs) provide an ample tool to capture syntax of many Indian languages. Tamil represents a special challenge to computational formalisms as it has extensive agglutinative morphology and a compara…

SentenceTAG