SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
Although the community has tackled the acquisition of high-quality Arabic pretraining data, we still lack large-scale, multi-turn Arabic datasets that include reasoning and tool calling. Naive translation can work at the pretraining scale, but post-training demands much higher quality, which requires a stricter approach to dataset curation. In this work, we introduce SmolKalam, a translation of Smoltalk2 that uses a multi-model ensemble translation pipeline, applies quality filtering, and examines effective translation techniques for traditional decoder-only models through ablations.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Improving Low-Resource Neural Machine Translation with Filtered Pseudo-Parallel Corpus
Large-scale parallel corpora are indispensable to train highly accurate machine translators. However, manually constructed large-scale parallel corpora are not freely available in many language pairs. In previous studies…
Language ModelingLanguage ModellingLow Resource Neural Machine TranslationLow-Resource Neural Machine Translation+3KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum Training
We present KinyaEmbed, the first dedicated sentence embedding model for Kinyarwanda, a morphologically rich Bantu language spoken by over 12 million people in Rwanda. Existing multilingual embedding models such as LaBSE,…
GFST: Gender-Filtered Self-Training for More Accurate Gender in Translation
Targeted evaluations have found that machine translation systems often output incorrect gender in translations, even when the gender is clear from context. Furthermore, these incorrectly gendered translations have the po…
Machine TranslationTranslationThe MLLP-UPV German-English Machine Translation System for WMT18
This paper describes the statistical machine translation system built by the MLLP research group of Universitat Polit{\`e}cnica de Val{\`e}ncia for the German→English news translation shared task of the EMNLP 2018 Thir…
Data AugmentationMachine TranslationTranslationImproving NMT via Filtered Back Translation
Document-Level Machine Translation (MT) has become an active research area among the NLP community in recent years. Unlike sentence-level MT, which translates the sentences independently, document-level MT aims to utiliz…
Document Level Machine TranslationMachine TranslationNMTSentence+1