paper-with-me

홈 › Papers

Syntactic Inductive Bias in Transformer Language Models: Especially Helpful for Low-Resource Languages?

2023-11-01 · Luke Gessler, Nathan Schneider

A line of work on Transformer-based language models such as BERT has attempted to use syntactic inductive bias to enhance the pretraining process, on the theory that building syntactic structure into the training process should reduce the amount of data needed for training. But such methods are often tested for high-resource languages such as English. In this work, we investigate whether these methods can compensate for data sparseness in low-resource languages, hypothesizing that they ought to be more effective for low-resource languages. We experiment with five low-resource languages: Uyghur, Wolof, Maltese, Coptic, and Ancient Greek. We find that these syntactic inductive bias methods produce uneven results in low-resource settings, and provide surprisingly little benefit in most cases.

📄 PDF Abstract BibTeX arXiv:2311.00268

Code (1)

lgessler/lr-sib 공식 구현 pytorch

Tasks

Inductive Bias

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Weight Decay 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.

Similar Papers 제목 키워드 기반

Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations

2024-07-05 · Matthias Lindemann, Alexander Koller, Ivan Titov

Models need appropriate inductive biases to effectively learn from small amounts of data and generalize systematically outside of the training distribution. While Transformers are highly versatile and powerful, they can …

ChunkingFew-Shot LearningInductive BiasSemantic Parsing

Sneaking Syntax into Transformer Language Models with Tree Regularization

2024-11-28 · Ananjan Nandi, Christopher D. Manning, Shikhar Murty

While compositional accounts of human language understanding are based on a hierarchical tree-like process, neural models like transformers lack a direct inductive bias for such tree structures. Introducing syntactic ind…

Inductive Bias

How to Plant Trees in Language Models: Data and Architectural Effects on the Emergence of Syntactic Inductive Biases

2023-05-31 · Aaron Mueller, Tal Linzen

Accurate syntactic representations are essential for robust generalization in natural language. Recent work has found that pre-training can teach language models to rely on hierarchical syntactic features - as opposed to…

DecoderInductive BiasLanguage Acquisition

Syntactic Inductive Biases for Deep Learning Methods

2022-06-08 · Yikang Shen

In this thesis, we try to build a connection between the two schools by introducing syntactic inductive biases for deep learning models. We propose two families of inductive biases, one for constituency structure and ano…

Deep LearningInductive Bias

Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language Models

2021-09-22 · EMNLP 2021 11 · Yiwen Wang, Jennifer Hu, Roger Levy, Peng Qian

Prior work has shown that structural supervision helps English language models learn generalizations about syntactic phenomena such as subject-verb agreement. However, it remains unclear if such an inductive bias would a…

Inductive Bias