paper-with-me

Papers

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

2026-06-29 · M. K. Arabov arxiv

Russian and Arabic are among the major languages of scientific communication. Language barriers impede the exchange of research results between these communities, which affects international collaboration and the progress of sustainability-related research. We present a benchmark for Arabic--Russian scientific translation. The benchmark includes a hybrid parallel corpus of about 27,000 sentence pairs, compiled from scientific abstracts and general-domain texts (religion, news, conversations). We fine-tune three multilingual language models -- mT5-base (580M parameters), NLLB-200-distilled-1.3B (1.3B), and Qwen2.5-7B-Instruct (7B) -- using LoRA with ranks 8, 16, 32, and 64. The Qwen2.5-7B model with QLoRA (rank 8) yields BLEU 23.15, chrF 43.89, BERTScore 0.906, and COMET 0.758. These are +4.36 BLEU and +0.051 COMET above the zero-shot baseline. Few-shot prompting with three examples does not improve performance, indicating that domain-specific fine-tuning is required. We release the models, the corpus, and the evaluation code. By lowering the language barrier for scientific texts, the work enables knowledge exchange between Arabic-speaking and Russian-speaking researchers. It contributes to sustainable partnerships (UN SDG 17) and innovation infrastructure (SDG 9), aligning with the conference's focus on technology-driven sustainable development.

📄 PDF Abstract BibTeX arXiv:2606.30943

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding

2026-01-27 · Shubham Patle, Sara Ghaboura, Hania Tariq, Mohammad Usman Khan 외 arxiv

Arabic calligraphy represents one of the richest visual traditions of the Arabic language, blending linguistic meaning with artistic form. Although multimodal models have advanced across languages, their ability to proce…

ASCAT: An Arabic Scientific Corpus and Benchmark for Advanced Translation Evaluation

2026-03-10 · Serry Sibaee, Khloud Al Jallad, Zineb Yousfi, Israa Elsayed Elhosiny 외 arxiv

We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine t…

Mubeen AI: A Specialized Arabic Language Model for Heritage Preservation and User Intent Understanding

2025-10-27 · Mohammed Aljafari, Ismail Alturki, Ahmed Mori, Yehya Kadumi arxiv

Mubeen is a proprietary Arabic language model developed by MASARAT SA, optimized for deep understanding of Arabic linguistics, Islamic studies, and cultural heritage. Trained on an extensive collection of authentic Arabi…

General KnowledgeIntent Detection

The United Nations Parallel Corpus v1.0

2016-05-01 · LREC 2016 5 · Micha{\l} Ziemski, Marcin Junczys-Dowmunt, Bruno Pouliquen

This paper describes the creation process and statistics of the official United Nations Parallel Corpus, the first parallel corpus composed from United Nations documents published by the original data creator. The parall…

Translation

The Multilingual Paraphrase Database

2014-05-01 · LREC 2014 5 · Juri Ganitkevitch, Chris Callison-Burch

We release a massive expansion of the paraphrase database (PPDB) that now includes a collection of paraphrases in 23 different languages. The resource is derived from large volumes of bilingual parallel data. Our collect…

Document SummarizationInformation RetrievalMachine TranslationMulti-Document Summarization+4