Learning curves theory for hierarchically compositional data with power-law distributed features
Recent theories suggest that Neural Scaling Laws arise whenever the task is linearly decomposed into power-law distributed units. Alternatively, scaling laws also emerge when data exhibit a hierarchically compositional structure, as is thought to occur in language and images. To unify these views, we consider classification and next-token prediction tasks based on probabilistic context-free grammars -- probabilistic models that generate data via a hierarchy of production rules. For classification, we show that having power-law distributed production rules results in a power-law learning curve with an exponent depending on the rules' distribution and a large multiplicative constant that depends on the hierarchical structure. By contrast, for next-token prediction, the distribution of production rules controls the local details of the learning curve, but not the exponent describing the large-scale behaviour.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Hierarchically Compositional Tasks and Deep Convolutional Networks
The main success stories of deep learning, starting with ImageNet, depend on deep convolutional networks, which on certain tasks perform significantly better than traditional shallow classifiers, such as support vector m…
Object RecognitionCompositional and Equilibrium-Free Conditions for Power System Stability -- Part I: Theory
Traditional centralized stability analysis struggles with scalability in large complex modern power grids. This two-part paper proposes a compositional and equilibrium-free approach to analyzing power system stability. I…
Compositional and Equilibrium-Free Conditions for Power System Stability -- Part II: Method and Application
This two-part paper proposes a compositional and equilibrium-free approach to analyzing power system stability. In Part I, we have established the stability theory and proposed stability conditions based on the delta dis…
Distributed ComputingEvaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing
Despite their strong performance on many tasks, pre-trained language models have been shown to struggle on out-of-distribution compositional generalization. Meanwhile, recent work has shown considerable improvements on m…
DecoderIn-Context LearningLanguage ModellingSemantic Parsing+1Learning curves for Gaussian process regression with power-law priors and targets
We characterize the power-law asymptotics of learning curves for Gaussian process regression (GPR) under the assumption that the eigenspectrum of the prior and the eigenexpansion coefficients of the target function follo…
GPRregression