Domain Adaptation for Arabic Machine Translation: The Case of Financial Texts
Neural machine translation (NMT) has shown impressive performance when trained on large-scale corpora. However, generic NMT systems have demonstrated poor performance on out-of-domain translation. To mitigate this issue, several domain adaptation methods have recently been proposed which often lead to better translation quality than genetic NMT systems. While there has been some continuous progress in NMT for English and other European languages, domain adaption in Arabic has received little attention in the literature. The current study, therefore, aims to explore the effectiveness of domain-specific adaptation for Arabic MT (AMT), in yet unexplored domain, financial news articles. To this end, we developed carefully a parallel corpus for Arabic-English (AR- EN) translation in the financial domain for benchmarking different domain adaptation methods. We then fine-tuned several pre-trained NMT and Large Language models including ChatGPT-3.5 Turbo on our dataset. The results showed that the fine-tuning is successful using just a few well-aligned in-domain AR-EN segments. The quality of ChatGPT translation was superior than other models based on automatic and human evaluations. To the best of our knowledge, this is the first work on fine-tuning ChatGPT towards financial domain transfer learning. To contribute to research in domain translation, we made our datasets and fine-tuned models available at https://huggingface.co/asas-ai/.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesBenchmarkingDomain AdaptationMachine TranslationNMTTransfer LearningTranslationSimilar Papers 제목 키워드 기반
Domain and Dialect Adaptation for Machine Translation into Egyptian Arabic
Domain-Specific Text Generation for Machine Translation
Preservation of domain knowledge from the source to target is crucial in any translation workflow. It is common in the translation industry to receive highly specialized projects, where there is hardly any parallel in-do…
Data AugmentationDomain AdaptationMachine TranslationPrompt Engineering+2QCRI’s Machine Translation Systems for IWSLT’16
This paper describes QCRI’s machine translation systems for the IWSLT 2016 evaluation campaign. We participated in the Arabic→English and English→Arabic tracks. We built both Phrase-based and Neural machine translation m…
Domain AdaptationLanguage ModelingLanguage ModellingMachine Translation+2QCRI Machine Translation Systems for IWSLT 16
This paper describes QCRI's machine translation systems for the IWSLT 2016 evaluation campaign. We participated in the Arabic->English and English->Arabic tracks. We built both Phrase-based and Neural machine translation…
Domain AdaptationLanguage ModelingLanguage ModellingMachine Translation+2Incremental Domain Adaptation for Neural Machine Translation in Low-Resource Settings
We study the problem of incremental domain adaptation of a generic neural machine translation model with limited resources (e.g., budget and time) for human translations or model training. In this paper, we propose a nov…
Active LearningDomain AdaptationInformativenessMachine Translation+3