paper-with-me

Papers

SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data

2025-11-23 · Sultan Alrashed, Chadi Helwe, Francesco Orabona arxiv

Although the community has tackled the acquisition of high-quality Arabic pretraining data, we still lack large-scale, multi-turn Arabic datasets that include reasoning and tool calling. Naive translation can work at the pretraining scale, but post-training demands much higher quality, which requires a stricter approach to dataset curation. In this work, we introduce SmolKalam, a translation of Smoltalk2 that uses a multi-model ensemble translation pipeline, applies quality filtering, and examines effective translation techniques for traditional decoder-only models through ablations.

📄 PDF Abstract BibTeX arXiv:2511.18411

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Improving Low-Resource Neural Machine Translation with Filtered Pseudo-Parallel Corpus

2017-11-01 · WS 2017 11 · Aizhan Imankulova, Takayuki Sato, Mamoru Komachi

Large-scale parallel corpora are indispensable to train highly accurate machine translators. However, manually constructed large-scale parallel corpora are not freely available in many language pairs. In previous studies…

Language ModelingLanguage ModellingLow Resource Neural Machine TranslationLow-Resource Neural Machine Translation+3

KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum Training

2026-08-27 · Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre arxiv

We present KinyaEmbed, the first dedicated sentence embedding model for Kinyarwanda, a morphologically rich Bantu language spoken by over 12 million people in Rwanda. Existing multilingual embedding models such as LaBSE,…

GFST: Gender-Filtered Self-Training for More Accurate Gender in Translation

2021-11-01 · EMNLP 2021 11 · Prafulla Kumar Choubey, Anna Currey, Prashant Mathur, Georgiana Dinu

Targeted evaluations have found that machine translation systems often output incorrect gender in translations, even when the gender is clear from context. Furthermore, these incorrectly gendered translations have the po…

Machine TranslationTranslation

The MLLP-UPV German-English Machine Translation System for WMT18

2018-10-01 · WS 2018 10 · Javier Iranzo-S{\'a}nchez, Pau Baquero-Arnal, Gon{\c{c}}al V. Garc{\'e}s D{\'\i}az-Mun{\'\i}o, Adri{\`a} Mart{\'\i}nez-Villaronga 외

This paper describes the statistical machine translation system built by the MLLP research group of Universitat Polit{\`e}cnica de Val{\`e}ncia for the German→English news translation shared task of the EMNLP 2018 Thir…

Data AugmentationMachine TranslationTranslation

Improving NMT via Filtered Back Translation

2020-12-01 · AACL (WAT) 2020 12 · Nikhil Jaiswal, Mayur Patidar, Surabhi Kumari, Manasi Patwardhan 외

Document-Level Machine Translation (MT) has become an active research area among the NLP community in recent years. Unlike sentence-level MT, which translates the sentences independently, document-level MT aims to utiliz…

Document Level Machine TranslationMachine TranslationNMTSentence+1