paper-with-me

홈 › Papers

Trained on 100 million words and still in shape: BERT meets British National Corpus

2023-03-17 · David Samuel, Andrey Kutuzov, Lilja Øvrelid, Erik Velldal

While modern masked language models (LMs) are trained on ever larger corpora, we here explore the effects of down-scaling training to a modestly-sized but representative, well-balanced, and publicly available English text source -- the British National Corpus. We show that pre-training on this carefully curated corpus can reach better performance than the original BERT model. We argue that this type of corpora has great potential as a language modeling benchmark. To showcase this potential, we present fair, reproducible and data-efficient comparative studies of LMs, in which we evaluate several training objectives and model architectures and replicate previous empirical results in a systematic way. We propose an optimized LM architecture called LTG-BERT.

📄 PDF Abstract BibTeX arXiv:2303.09859

Code (2)

ltgoslo/ltg-bert 공식 구현 pytorch
ltgoslo/factorizer pytorch

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

SimpleBERT: A Pre-trained Model That Learns to Generate Simple Words

2022-04-16 · Renliang Sun, Xiaojun Wan

Pre-trained models are widely used in the tasks of natural language processing nowadays. However, in the specific field of text simplification, the research on improving pre-trained models is still blank. In this work, w…

Language ModelingLanguage ModellingLexical SimplificationMasked Language Modeling+2

PhayaThaiBERT: Enhancing a Pretrained Thai Language Model with Unassimilated Loanwords

2023-11-21 · Panyut Sriwirote, Jalinee Thapiang, Vasan Timtong, Attapol T. Rutherford

While WangchanBERTa has become the de facto standard in transformer-based Thai language modeling, it still has shortcomings in regard to the understanding of foreign words, most notably English words, which are often bor…

Language ModelingLanguage Modelling

Pretraining without Wordpieces: Learning Over a Vocabulary of Millions of Words

2022-02-24 · Zhangyin Feng, Duyu Tang, Cong Zhou, Junwei Liao 외

The standard BERT adopts subword-based tokenization, which may break a word into two or more wordpieces (e.g., converting "lossless" to "loss" and "less"). This will bring inconvenience in following situations: (1) what …

ChunkingCloze TestMachine Reading ComprehensionNatural Language Understanding+4

WhisBERT: Multimodal Text-Audio Language Modeling on 100M Words

2023-12-05 · Lukas Wolf, Greta Tuckute, Klemen Kotar, Eghbal Hosseini 외

Training on multiple modalities of input can augment the capabilities of a language model. Here, we ask whether such a training regime can improve the quality and efficiency of these systems as well. We focus on text--au…

Language ModelingLanguage Modelling

AMBERT: A Pre-trained Language Model with Multi-Grained Tokenization

2020-08-27 · Findings (ACL) 2021 8 · Xinsong Zhang, Pengshuai Li, Hang Li

Pre-trained language models such as BERT have exhibited remarkable performances in many tasks in natural language understanding (NLU). The tokens in the models are usually fine-grained in the sense that for languages lik…

Language ModelingLanguage ModellingNatural Language Understanding