paper-with-me

홈 › Papers

AI-Driven Generation of Old English: A Framework for Low-Resource Languages

2025-07-27 · Rodrigo Gabriel Salazar Alva, Matías Nuñez, Cristian López, Javier Martín Arista arxiv

Preserving ancient languages is essential for understanding humanity's cultural and linguistic heritage, yet Old English remains critically under-resourced, limiting its accessibility to modern natural language processing (NLP) techniques. We present a scalable framework that uses advanced large language models (LLMs) to generate high-quality Old English texts, addressing this gap. Our approach combines parameter-efficient fine-tuning (Low-Rank Adaptation, LoRA), data augmentation via backtranslation, and a dual-agent pipeline that separates the tasks of content generation (in English) and translation (into Old English). Evaluation with automated metrics (BLEU, METEOR, and CHRF) shows significant improvements over baseline models, with BLEU scores increasing from 26 to over 65 for English-to-Old English translation. Expert human assessment also confirms high grammatical accuracy and stylistic fidelity in the generated texts. Beyond expanding the Old English corpus, our method offers a practical blueprint for revitalizing other endangered languages, effectively uniting AI innovation with the goals of cultural preservation.

📄 PDF Abstract BibTeX arXiv:2507.20111

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningData Augmentation

Similar Papers 제목 키워드 기반

Language Agnostic Data-Driven Inverse Text Normalization

2023-01-20 · Szu-Jui Chen, Debjyoti Paul, Yutong Pang, Peng Su 외

With the emergence of automatic speech recognition (ASR) models, converting the spoken form text (from ASR) to the written form is in urgent need. This inverse text normalization (ITN) problem attracts the attention of r…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationForm+3

Aligning LLMs for Multilingual Consistency in Enterprise Applications

2025-09-28 · Amit Agarwal, Hansa Meghwani, Hitesh Laxmichand Patel, Tao Sheng 외 arxiv

Large language models (LLMs) remain unreliable for global enterprise applications due to substantial performance gaps between high-resource and mid/low-resource languages, driven by English-centric pretraining and intern…

Information Retrieval

XAlign: Cross-lingual Fact-to-Text Alignment and Generation for Low-Resource Languages

2022-02-01 · Tushar Abhishek, Shivprasad Sagare, Bhavyajeet Singh, Anubhav Sharma 외

Multiple critical scenarios (like Wikipedia text generation given English Infoboxes) need automated generation of descriptive text in low resource (LR) languages from English fact triples. Previous work has focused on En…

Data-to-Text GenerationDescriptiveText Generation

Data-to-text Generation for Severely Under-Resourced Languages with GPT-3.5: A Bit of Help Needed from Google Translate

2023-08-19 · Michela Lorandi, Anya Belz

LLMs like GPT are great at tasks involving English which dominates in their training data. In this paper, we look at how they cope with tasks involving languages that are severely under-represented in their training data…

Data-to-Text GenerationPrompt EngineeringText GenerationTranslation

Enhancing Multilingual RAG Systems with Debiased Language Preference-Guided Query Fusion

2026-01-06 · Jeonghyun Park, Byeongjeong Kim, Seojin Hwang, Hwanhee Lee arxiv

Multilingual Retrieval-Augmented Generation (mRAG) systems often exhibit a perceived preference for high-resource languages, particularly English, resulting in the widespread adoption of English pivoting. While prior stu…