Unraveling Syntax: Language Modeling and the Substructure of Grammars
While language models achieve impressive results, their learning dynamics are far from understood. Many domains of interest -- such as natural language syntax, coding languages, arithmetic -- are captured by context-free grammars (CFGs). In this work, we extend prior work on neural language modeling of CFGs in a novel direction: how language modeling behaves with respect to CFG substructure, namely subgrammars. We define subgrammars, and prove a set of fundamental theorems connecting language modeling and subgrammars. We show that language modeling loss recurses linearly over its top-level subgrammars; applied recursively, the loss decomposes into losses for "irreducible" subgrammars. Under additional assumptions, and empirically, parametrized models learn subgrammars in parallel, unlike children who first master simple substructures. We find that subgrammar pretraining can improve final performance, but only for tiny models relative to the grammar, while alignment analyses show that pretraining consistently leads to internal representations that better reflect the grammar's substructure.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Unsupervised Recurrent Neural Network Grammars
Recurrent neural network grammars (RNNG) are generative models of language which jointly model syntax and surface structure by incrementally generating a syntax tree and sentence in a top-down, left-to-right order. Super…
Constituency Grammar InductionLanguage ModelingLanguage ModellingSentence+1What Do Recurrent Neural Network Grammars Learn About Syntax?
Recurrent neural network grammars (RNNG) are a recently proposed probabilistic generative modeling family for natural language. They show state-of-the-art language modeling and parsing performance. We investigate what in…
Constituency ParsingDependency ParsingLanguage ModelingLanguage ModellingProbabilistic modeling of the syntax semantic interface using probabilistic context free grammars, Experiments with FrameNet (Mod\'elisation probabiliste de l'interface syntaxe s\'emantique \`a l'aide de grammaires hors contexte probabilistes Exp\'eriences avec FrameNet) [in French]
Transformer Grammars: Augmenting Transformer Language Models with Syntactic Inductive Biases at Scale
We introduce Transformer Grammars (TGs), a novel class of Transformer language models that combine (i) the expressive power, scalability, and strong performance of Transformers and (ii) recursive syntactic compositions, …
Inductive BiasLanguage ModelingLanguage ModellingSentenceDependency Transformer Grammars: Integrating Dependency Structures into Transformer Language Models
Syntactic Transformer language models aim to achieve better generalization through simultaneously modeling syntax trees and sentences. While prior work has been focusing on adding constituency-based structures to Transfo…
ARCInductive BiasLanguage ModelingLanguage Modelling