Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models
As large language models (LLMs) grow and develop, so do their data demands. This is especially true for multilingual LLMs, where the scarcity of high-quality and readily available data online has led to a multitude of synthetic dataset generation approaches. A key technique in this space is machine translation (MT), where high-quality English text is adapted to a target, comparatively low-resource language. This report introduces FineWeb-Edu-Ar, a machine-translated version of the exceedingly popular (deduplicated) FineWeb-Edu dataset from HuggingFace. To the best of our knowledge, FineWeb-Edu-Ar is the largest publicly available machine-translated Arabic dataset out there, with its size of 202B tokens of an Arabic-trained tokenizer.
Code (0)
등록된 구현이 없습니다.
Tasks
Dataset GenerationMachine TranslationSimilar Papers 제목 키워드 기반
Simple Automatic Post-editing for Arabic-Japanese Machine Translation
A common bottleneck for developing machine translation (MT) systems for some language pairs is the lack of direct parallel translation data sets, in general and in certain domains. Alternative solutions such as zero-shot…
ArticlesAutomatic Post-EditingMachine TranslationTranslationBuilding an Arabic Machine Translation Post-Edited Corpus: Guidelines and Annotation
We present our guidelines and annotation procedure to create a human corrected machine translated post-edited corpus for the Modern Standard Arabic. Our overarching goal is to use the annotated corpus to develop automati…
ArticlesMachine TranslationTranslationMultilingual Language Model Pretraining using Machine-translated Data
High-resource languages such as English, enables the pretraining of high-quality large language models (LLMs). The same can not be said for most other languages as LLMs still underperform for non-English languages, likel…
Language ModelingLanguage ModellingA Parallel Corpus for Evaluating Machine Translation between Arabic and European Languages
We present Arab-Acquis, a large publicly available dataset for evaluating machine translation between 22 European languages and Arabic. Arab-Acquis consists of over 12,000 sentences from the JRC-Acquis (Acquis Communauta…
BenchmarkingMachine TranslationTranslationASCAT: An Arabic Scientific Corpus and Benchmark for Advanced Translation Evaluation
We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine t…