LuxemBERT: Simple and Practical Data Augmentation in Language Model Pre-Training for Luxembourgish
Pre-trained Language Models such as BERT have become ubiquitous in NLP where they have achieved state-of-the-art performance in most NLP tasks. While these models are readily available for English and other widely spoken languages, they remain scarce for low-resource languages such as Luxembourgish. In this paper, we present LuxemBERT, a BERT model for the Luxembourgish language that we create using the following approach: we augment the pre-training dataset by considering text data from a closely related language that we partially translate using a simple and straightforward method. We are then able to produce the LuxemBERT model, which we show to be effective for various NLP tasks: it outperforms a simple baseline built with the available Luxembourgish text data as well the multilingual mBERT model, which is currently the only option for transformer-based language models in Luxembourgish. Furthermore, we present datasets for various downstream NLP tasks that we created for this study and will make available to researchers on request.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
MulDA: A Multilingual Data Augmentation Framework for Low-Resource Cross-Lingual NER
Named Entity Recognition (NER) for low-resource languages is a both practical and challenging research problem. This paper addresses zero-shot transfer for cross-lingual NER, especially when the amount of source-language…
Cross-Lingual NERCross-Lingual TransferData AugmentationDiversity+7An Analysis of Simple Data Augmentation for Named Entity Recognition
Simple yet effective data augmentation techniques have been proposed for sentence-level and sentence-pair natural language processing tasks. Inspired by these efforts, we design and compare data augmentation for named en…
Data Augmentationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1From Language Models to Practical Self-Improving Computer Agents
We develop a simple and straightforward methodology to create AI computer agents that can carry out diverse computer tasks and self-improve by developing tools and augmentations to enable themselves to solve increasingly…
Prompt EngineeringRetrievalLidarAugment: Searching for Scalable 3D LiDAR Data Augmentations
Data augmentations are important in training high-performance 3D object detectors for point clouds. Despite recent efforts on designing new data augmentations, perhaps surprisingly, most state-of-the-art 3D detectors onl…
3D Object DetectionData Augmentationobject-detectionObject DetectionCheap and Good? Simple and Effective Data Augmentation for Low Resource Machine Reading
We propose a simple and effective strategy for data augmentation for low-resource machine reading comprehension (MRC). Our approach first pretrains the answer extraction components of a MRC system on the augmented data t…
Data AugmentationMachine Reading ComprehensionReading ComprehensionRetrieval