paper-with-me

홈 › Papers

Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT

2025-10-09 · Noor Ul Zain, Mohsin Raza, Ahsan Adeel arxiv

We show that a tiny Co$^4$ machine(Adeel,2025) with a single layer, two heads, and 8M parameters, operating at an approximate cost of $O(N)$ (where $N$ is the number of input tokens), outpaces the BabyLM Challenge baselines GPT-2 (124M, 12 layers, $O(N^2))$ and GPT-BERT (30M, 12 layers, $O(N^2))$ in just two epochs, while both are trained for ten. Co$^4$ achieves orders-of-magnitude greater training efficiency on 10M tokens, demonstrating highly sample efficient pretraining. Using the BabyLM challenge evaluation pipeline across complex benchmarks, Co$^4$ exhibits strong zero-shot and fine-tuning performance on SuperGLUE tasks. Specifically, Co$^4$ outperforms GPT-2 on 5 out of 7 zero-shot metrics and 6 out of 7 fine-tuning tasks, and GPT-BERT on 4 out of 7 metrics in both cases. These results suggest the need to rethink prevailing deep learning paradigms and associated scaling laws.

📄 PDF Abstract BibTeX arXiv:2510.08404

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TinyBERT: Distilling BERT for Natural Language Understanding

2019-09-23 · Findings of the Association for Computational Linguistics 2020 · Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang 외

Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is diffi…

Knowledge DistillationLanguage ModellingLinguistic AcceptabilityNatural Language Inference+5

Dynamic-TinyBERT: Boost TinyBERT's Inference Efficiency by Dynamic Sequence Length

2021-11-18 · Shira Guskin, Moshe Wasserblat, Ke Ding, Gyuwan Kim

Limited computational budgets often prevent transformers from being used in production and from having their high accuracy utilized. TinyBERT addresses the computational efficiency by self-distilling BERT into a smaller …

Computational EfficiencyHyperparameter OptimizationQuestion Answering

EELBERT: Tiny Models through Dynamic Embeddings

2023-10-31 · Gabrielle Cohn, Rishika Agarwal, Deepanshu Gupta, Siddharth Patwardhan

We introduce EELBERT, an approach for compression of transformer-based models (e.g., BERT), with minimal impact on the accuracy of downstream tasks. This is achieved by replacing the input embedding layer of the model wi…

Investigating Mixture of Experts in Dense Retrieval

2024-12-16 · Effrosyni Sokli, Pranav Kasela, Georgios Peikos, Gabriella Pasi

While Dense Retrieval Models (DRMs) have advanced Information Retrieval (IR), one limitation of these neural models is their narrow generalizability and robustness. To cope with this issue, one can leverage the Mixture-o…

Information RetrievalMixture-of-ExpertsRetrieval

AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models

2021-07-29 · ACL 2021 5 · Yichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang 외

Pre-trained language models (PLMs) have achieved great success in natural language processing. Most of PLMs follow the default setting of architecture hyper-parameters (e.g., the hidden dimension is a quarter of the inte…

Neural Architecture SearchOne-Shot Learning