paper-with-me

홈 › Papers

BAMBINO-LM: (Bilingual-)Human-Inspired Continual Pretraining of BabyLM

2024-06-17 · Zhewen Shen, Aditya Joshi, Ruey-Cheng Chen

Children from bilingual backgrounds benefit from interactions with parents and teachers to re-acquire their heritage language. In this paper, we investigate how this insight from behavioral study can be incorporated into the learning of small-scale language models. We introduce BAMBINO-LM, a continual pre-training strategy for BabyLM that uses a novel combination of alternation and PPO-based perplexity reward induced from a parent Italian model. Upon evaluation on zero-shot classification tasks for English and Italian, BAMBINO-LM improves the Italian language capability of a BabyLM baseline. Our ablation analysis demonstrates that employing both the alternation strategy and PPO-based modeling is key to this effectiveness gain. We also show that, as a side effect, the proposed method leads to a similar degradation in L1 effectiveness as human children would have had in an equivalent learning scenario. Through its modeling and findings, BAMBINO-LM makes a focused contribution to the pre-training of small-scale language models by first developing a human-inspired strategy for pre-training and then showing that it results in behaviours similar to that of humans.

📄 PDF Abstract BibTeX arXiv:2406.11418

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Pretrainingzero-shot-classificationZero-Shot Learning

Similar Papers 제목 키워드 기반

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

2025-05-31 · Shaoxiong Ji, Zihao Li, Jaakko Paavola, Indraneil Paul 외

This paper investigates a critical design decision in the practice of massively multilingual continual pre-training -- the inclusion of parallel data. Specifically, we study the impact of bilingual translation data for m…

Translation

Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi

2026-03-03 · Shiza Fatimah, Aniket Sen, Sophia Falk, Florian Mai 외 arxiv

The dominance of large multilingual foundation models has widened linguistic inequalities in Natural Language Processing (NLP), often leaving low-resource languages underrepresented. This paper introduces LilMoo, a 0.6-b…

Continual Pretraining

A Study on Efficiency in Continual Learning Inspired by Human Learning

2020-10-28 · Philip J. Ball, Yingzhen Li, Angus Lamb, Cheng Zhang

Humans are efficient continual learning systems; we continually learn new skills from birth with finite cells and resources. Our learning is highly optimized both in terms of capacity and time while not suffering from ca…

Continual Learning

Biologically-Inspired Continual Learning of Human Motion Sequences

2022-11-02 · Joachim Ott, Shih-Chii Liu

This work proposes a model for continual learning on tasks involving temporal sequences, specifically, human motions. It improves on a recently proposed brain-inspired replay model (BI-R) by building a biologically-inspi…

Continual LearningTemporal Sequences

The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining

2026-01-01 · Jiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani 외 arxiv

Multilingual large language models achieve impressive cross-lingual performance despite largely monolingual pretraining. While bilingual data in pretraining corpora is widely believed to enable these abilities, details o…