paper-with-me

홈 › Papers

Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models

2026-05-18 · Spyridon Mavromatis, Sokratis Sofianopoulos, Prokopis Prokopidis, Maria Giagkou arxiv

Machine Translation (MT) for Ancient Greek (AG) to Modern Greek (MG) is a low-resource task, constrained by the lack of large-scale, high-quality parallel data. We address this gap by introducing the AG-MG Parallel Corpus, a new resource containing 132,481 sentence-aligned pairs derived from literary, historical, and biblical texts. We present a novel corpus creation pipeline that combines web-scraped, excerpt-level data with a multi-stage sentence-level alignment, and refinement process. Our method uses VecAlign with LaBSE embeddings, which we first fine-tune on a manually-aligned AG-MG subset, followed by an LLM-based error/misalignment correction phase using Gemini 2.5 Flash to ensure high alignment quality. Furthermore, we provide the first comprehensive benchmark of modern MT models on this task, evaluating three fine-tuning strategies across NMT models (NLLB, M2M100) and a Greek LLM (Llama-Krikri-8B). Our experiments show that fine-tuning yields significant improvements over base models, increasing performance by up to +10.3 BLEU points. Specifically, full-parameter fine-tuning of Llama-Krikri-8B achieves the highest overall performance with a BLEU score of 13.16, while the QLoRA-adapted M2M100-1.2B model demonstrates the largest relative gains and highly competitive results. Our dataset and models represent a significant contribution to Greek NLP.

📄 PDF Abstract BibTeX arXiv:2605.18504

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Ancient Greek WordNet Meets the Dynamic Lexicon: the Example of the Fragments of the Greek Historians

2016-01-01 · GWC 2016 1 · Monica Berti, Yuri Bizzoni, Federico Boschetti, Gregory R. Crane 외

The Ancient Greek WordNet (AGWN) and the Dynamic Lexicon (DL) are multilingual resources to study the lexicon of Ancient Greek texts and their translations. Both AGWN and DL are works in progress that need accuracy impro…

An automatic model and Gold Standard for translation alignment of Ancient Greek

2022-06-01 · LREC 2022 6 · Tariq Yousef, Chiara Palladino, Farnoosh Shamsian, Anise d’Orange Ferreira 외

This paper illustrates a workflow for developing and evaluating automatic translation alignment models for Ancient Greek. We designed an annotation Style Guide and a gold standard for the alignment of Ancient Greek-Engli…

Translation

Sentence Embedding Models for Ancient Greek Using Multilingual Knowledge Distillation

2023-08-24 · Kevin Krahn, Derrick Tate, Andrew C. Lamicela

Contextual language models have been trained on Classical languages, including Ancient Greek and Latin, for tasks such as lemmatization, morphological tagging, part of speech tagging, authorship attribution, and detectio…

Authorship AttributionKnowledge DistillationLemmatizationMorphological Tagging+10

A Pilot Study for BERT Language Modelling and Morphological Analysis for Ancient and Medieval Greek

2021-11-01 · EMNLP (LaTeCHCLfL, CLFL, LaTeCH) 2021 11 · Pranaydeep Singh, Gorik Rutten, Els Lefever

This paper presents a pilot study to automatic linguistic preprocessing of Ancient and Byzantine Greek, and morphological analysis more specifically. To this end, a novel subword-based BERT language model was trained on …

Language ModelingLanguage ModellingMorphological Analysis

Automatic Translation Alignment for Ancient Greek and Latin

2022-06-01 · LT4HALA (LREC) 2022 6 · Tariq Yousef, Chiara Palladino, David J. Wright, Monica Berti

This paper presents the results of automatic translation alignment experiments on a corpus of texts in Ancient Greek translated into Latin. We used a state-of-the-art alignment workflow based on a contextualized multilin…

Language ModelingLanguage ModellingTranslation