paper-with-me

Papers

ToddlerBERTa: Exploiting BabyBERTa for Grammar Learning and Language Understanding

2023-08-30 · Omer Veysel Cagatan

We present ToddlerBERTa, a BabyBERTa-like language model, exploring its capabilities through five different models with varied hyperparameters. Evaluating on BLiMP, SuperGLUE, MSGS, and a Supplement benchmark from the BabyLM challenge, we find that smaller models can excel in specific tasks, while larger models perform well with substantial data. Despite training on a smaller dataset, ToddlerBERTa demonstrates commendable performance, rivalling the state-of-the-art RoBERTa-base. The model showcases robust language understanding, even with single-sentence pretraining, and competes with baselines that leverage broader contextual information. Our work provides insights into hyperparameter choices, and data utilization, contributing to the advancement of language models.

📄 PDF Abstract BibTeX arXiv:2308.16336

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingSentence

Similar Papers 제목 키워드 기반

BabyBERTa: Learning More Grammar With Small-Scale Child-Directed Language

2021-11-01 · CoNLL (EMNLP) 2021 11 · Philip A. Huebner, Elior Sulem, Fisher Cynthia, Dan Roth

Transformer-based language models have taken the NLP world by storm. However, their potential for addressing important questions in language acquisition research has been largely ignored. In this work, we examined the gr…

Language Acquisition

Tolerance Principle and Small Language Model Learning

2026-01-17 · Adam E. Friedman, Stevan Harnad, Rushen Shi arxiv

Modern language models like GPT-3, BERT, and LLaMA require massive training data, yet with sufficient training they reliably learn to distinguish grammatical from ungrammatical sentences. Children aged as young as 14 mon…

On the effect of curriculum learning with developmental data for grammar acquisition

2023-10-31 · Mattia Opper, J. Morrison, N. Siddharth

This work explores the degree to which grammar acquisition is driven by language `simplicity' and the source modality (speech vs. text) of data. Using BabyBERTa as a probe, we find that grammar acquisition is largely dri…

Learning from Child-Directed Speech in Two-Language Scenarios: A French-English Case Study

2026-03-13 · Liel Binyamin, Elior Sulem arxiv

Research on developmentally plausible language models has largely focused on English, leaving open questions about multilingual settings. We present a systematic study of compact language models by extending BabyBERTa to…

Exploiting Language Variants Via Grammar Parsing Having Morphologically Rich Information

2014-10-01 · WS 2014 10 · Qaiser Abbas