paper-with-me

홈 › Papers

On the Inductive Bias of Masked Language Modeling: From Statistical to Syntactic Dependencies

2021-04-12 · NAACL 2021 4 · Tianyi Zhang, Tatsunori Hashimoto

We study how masking and predicting tokens in an unsupervised fashion can give rise to linguistic structures and downstream performance gains. Recent theories have suggested that pretrained language models acquire useful inductive biases through masks that implicitly act as cloze reductions for downstream tasks. While appealing, we show that the success of the random masking strategy used in practice cannot be explained by such cloze-like masks alone. We construct cloze-like masks using task-specific lexicons for three different classification datasets and show that the majority of pretrained performance gains come from generic masks that are not associated with the lexicon. To explain the empirical success of these generic masks, we demonstrate a correspondence between the Masked Language Model (MLM) objective and existing methods for learning statistical dependencies in graphical models. Using this, we derive a method for extracting these learned statistical dependencies in MLMs and show that these dependencies encode useful inductive biases in the form of syntactic structures. In an unsupervised parsing evaluation, simply forming a minimum spanning tree on the implied statistical dependence structure outperforms a classic method for unsupervised parsing (58.74 vs. 55.91 UUAS).

📄 PDF Abstract BibTeX arXiv:2104.05694

Code (1)

tatsu-lab/mlm_inductive_bias 공식 구현 pytorch

Tasks

Inductive BiasLanguage ModelingLanguage ModellingMasked Language Modeling

Similar Papers 제목 키워드 기반

MaDis-Stereo: Enhanced Stereo Matching via Distilled Masked Image Modeling

2024-09-04 · Jihye Ahn, Hyesong Choi, SooMin Kim, Dongbo Min

In stereo matching, CNNs have traditionally served as the predominant architectures. Although Transformer-based stereo models have been studied recently, their performance still lags behind CNN-based stereo models due to…

Depth EstimationDepth PredictionImage ReconstructionInductive Bias+1

Profile Prediction: An Alignment-Based Pre-Training Task for Protein Sequence Models

2020-12-01 · Pascal Sturmfels, Jesse Vig, Ali Madani, Nazneen Fatema Rajani

For protein sequence datasets, unlabeled data has greatly outpaced labeled data due to the high cost of wet-lab characterization. Recent deep-learning approaches to protein prediction have shown that pre-training on unla…

Language ModelingLanguage ModellingMasked Language ModelingOpen-Ended Question Answering

A Cognitive Regularizer for Language Modeling

2021-05-15 · ACL 2021 5 · Jason Wei, Clara Meister, Ryan Cotterell

The uniform information density (UID) hypothesis, which posits that speakers behaving optimally tend to distribute information uniformly across a linguistic signal, has gained traction in psycholinguistics as an explanat…

Inductive BiasLanguage ModelingLanguage Modelling

Universal linguistic inductive biases via meta-learning

2020-06-29 · R. Thomas McCoy, Erin Grant, Paul Smolensky, Thomas L. Griffiths 외

How do learners acquire languages from the limited data available to them? This process must involve some inductive biases - factors that affect how a learner generalizes - but it is unclear which inductive biases can ex…

Language AcquisitionMeta-Learning

HeceTokenizer: A Syllable-Based Tokenization Approach for Turkish Retrieval

2026-04-12 · Senol Gulgonul arxiv

HeceTokenizer is a syllable-based tokenizer for Turkish that exploits the deterministic six-pattern phonological structure of the language to construct a closed, out-of-vocabulary (OOV)-free vocabulary of approximately 8…