paper-with-me

홈 › Papers

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi

2026-06-27 · Manasi Waghe, Danish Chandargi, Mohammad Aamir Rayyan, Raviraj Joshi, A. R. Deshpande arxiv

Government documents in India are predominantly issued in regional languages such as Marathi, creating substantial accessibility barriers for non-native readers, interstate administrative bodies, and policy analysts. Although recent advances in neural machine translation have improved sentence-level translation quality, existing systems largely neglect document structure, formatting integrity, and domain-specific terminology, thereby limiting their applicability to official documentation. This paper presents a structure-preserving Marathi-to-English government document translation framework capable of performing end-to-end document transformation while maintaining layout fidelity. The proposed system integrates layout-aware optical character recognition, coordinate-based text extraction, large language model based translation, and structured document reconstruction through HTML representations. By enforcing spatial alignment constraints and preserving hierarchical document elements, the framework ensures structural consistency between the source and translated documents. Experimental evaluation on real-world Marathi government PDFs demonstrates improved structural preservation, translation coherence, and terminological consistency compared to conventional text-only translation pipelines. The proposed framework contributes toward scalable multilingual accessibility solutions for e-governance and administrative document processing.

📄 PDF Abstract BibTeX arXiv:2606.28796

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Discourse Graph Guided Document Translation with Large Language Models

2025-11-10 · Viet-Thanh Pham, Minghan Wang, Hao-Han Liao, Thuy-Trang Vu arxiv

Adapting large language models to full document translation remains challenging due to the difficulty of capturing long-range dependencies and preserving discourse coherence throughout extended texts. While recent agenti…

Machine Translation

Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation

2026-03-23 · Ireh Kim, Tesia Sker, Chanwoo Kim arxiv

In Machine Translation, Large Language Models (LLMs) have generally underperformed compared to conventional encoder-decoder systems and thus see limited adoption. However, LLMs excel at modeling contextual information, m…

Machine Translation

CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia

2026-06-18 · Josef Jon, Ondřej Bojar arxiv

We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English, with smaller portions of Vietnamese, R…

Machine Translation

BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation

2026-05-11 · Qi Yang, Xiangyao Ma, Xiao Wang, Hao Wang 외 arxiv

As global cross-lingual communication intensifies, language barriers in visually rich documents such as PDFs remain a practical bottleneck. Existing document translation pipelines face a tension between linguistic proces…

Document Sub-structure in Neural Machine Translation

2019-12-13 · LREC 2020 5 · Radina Dobreva, Jie zhou, Rachel Bawden

Current approaches to machine translation (MT) either translate sentences in isolation, disregarding the context they appear in, or model context at the level of the full document, without a notion of any internal struct…

ArticlesMachine TranslationSentenceTranslation