paper-with-me

홈 › Papers

BabyLlama-2: Ensemble-Distilled Models Consistently Outperform Teachers With Limited Data

2024-09-25 · Jean-Loup Tastet, Inar Timiryasov

We present BabyLlama-2, a 345 million parameter model distillation-pretrained from two teachers on a 10 million word corpus for the BabyLM competition. On BLiMP and SuperGLUE benchmarks, BabyLlama-2 outperforms baselines trained on both 10 and 100 million word datasets with the same data mix, as well as its teacher models. Through an extensive hyperparameter sweep, we demonstrate that the advantages of distillation cannot be attributed to suboptimal hyperparameter selection of the teachers. Our findings underscore the need for further investigation into distillation techniques, particularly in data-limited settings.

📄 PDF Abstract BibTeX arXiv:2409.17312

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Choosy Babies Need One Coach: Inducing Mode-Seeking Behavior in BabyLlama with Reverse KL Divergence

2024-10-29 · Shaozhen Shi, Yevgen Matusevych, Malvina Nissim

This study presents our submission to the Strict-Small Track of the 2nd BabyLM Challenge. We use a teacher-student distillation setup with the BabyLLaMa model (Timiryasov and Tastet, 2023) as a backbone. To make the stud…

Adaptive Group Robust Ensemble Knowledge Distillation

2024-11-22 · Patrik Kenfack, Ulrich Aïvodji, Samira Ebrahimi Kahou

Neural networks can learn spurious correlations in the data, often leading to performance disparity for underrepresented subgroups. Studies have demonstrated that the disparity is amplified when knowledge is distilled fr…

Knowledge Distillation

Distilling Tabular Foundation Models for Structured Health Data

2026-05-18 · Aditya Tanna, Nassim Bouarour, Mohamed Bouadi, Vinay Kumar Sankarapu 외 arxiv

Tabular foundation models (TFMs) achieve strong performance on health datasets, but their inference cost and infrastructure requirements limit practical use. We study whether their predictive behavior can be transferred …

Knowledge Distillation

Monitored Distillation for Positive Congruent Depth Completion

2022-03-30 · Tian Yu Liu, Parth Agrawal, Allison Chen, Byung-Woo Hong 외

We propose a method to infer a dense depth map from a single image, its calibration, and the associated sparse point cloud. In order to leverage existing models (teachers) that produce putative depth maps, we propose an …

Depth CompletionImage ReconstructionKnowledge DistillationModel Selection

Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty

2023-08-03 · Inar Timiryasov, Jean-Loup Tastet

We present our submission to the BabyLM challenge, whose goal was to improve the sample efficiency of language models. We trained an ensemble consisting of a GPT-2 and small LLaMA models on the developmentally-plausible,…

Knowledge Distillation