On the importance of pre-training data volume for compact language models
Recent advances in language modeling have led to computationally intensive and resource-demanding state-of-the-art models. In an effort towards sustainable practices, we study the impact of pre-training data volume on compact language models. Multiple BERT-based models are trained on gradually increasing amounts of French text. Through fine-tuning on the French Question Answering Dataset (FQuAD), we observe that well-performing models are obtained with as little as 100 MB of text. In addition, we show that past critically low amounts of pre-training data, an intermediate pre-training step on the task-specific corpus does not yield substantial improvements.
Code (0)
등록된 구현이 없습니다.
Tasks
FQuADLanguage ModelingLanguage ModellingQuestion AnsweringSimilar Papers 제목 키워드 기반
Training-Free Zero-Shot Anomaly Detection in 3D Brain MRI with 2D Foundation Models
Zero-shot anomaly detection (ZSAD) has gained increasing attention in medical imaging as a way to identify abnormalities without task-specific supervision, but most advances remain limited to 2D datasets. Extending ZSAD …
Anomaly DetectionTrust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
Vision-language models (VLMs) extend the conventional large language models by integrating visual data, enabling richer multimodal reasoning and significantly broadens the practical applications of AI. However, including…
Multimodal ReasoningCost Volume Pyramid Based Depth Inference for Multi-View Stereo
We propose a cost volume-based neural network for depth inference from multi-view images. We demonstrate that building a cost volume pyramid in a coarse-to-fine manner instead of constructing a cost volume at a fixed res…
3D ReconstructionPoint CloudsRate-aware Compression for NeRF-based Volumetric Video
The neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage and transmission. To address these proble…
NeRFQuantizationInteractive Volume Visualization via Multi-Resolution Hash Encoding based Neural Representation
Neural networks have shown great potential in compressing volume data for visualization. However, due to the high cost of training and inference, such volumetric neural representations have thus far only been applied to …
GPU