paper-with-me

홈 › Papers

Assessing and Improving Punctuation Robustness in English-Marathi Machine Translation

2025-12-28 · Kaustubh Shivshankar Shejole, Sourabh Deoghare, Pushpak Bhattacharyya arxiv

Neural Machine Translation (NMT) systems rely heavily on explicit punctuation cues to resolve semantic ambiguities in a source sentence. Inputting user-generated sentences, which are likely to contain missing or incorrect punctuation, results in fluent but semantically disastrous translations. This work attempts to highlight and address the problem of punctuation robustness of NMT systems through an English-to-Marathi translation. First, we introduce \textbf{\textit{Viram}}, a human-curated diagnostic benchmark of 54 punctuation-ambiguous English-Marathi sentence pairs to stress-test existing NMT systems. Second, we evaluate two simple remediation strategies: cascade-based \textit{restore-then-translate} and \textit{direct fine-tuning}. Our experimental results and analysis demonstrate that both strategies yield substantial NMT performance improvements. Furthermore, we find that current Large Language Models (LLMs) exhibit relatively poorer robustness in translating such sentences than these task-specific strategies, thus necessitating further research in this area. The code and dataset are available at https://github.com/KaustubhShejole/Viram_Marathi.

📄 PDF Abstract BibTeX arXiv:2601.09725

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

English-Marathi Neural Machine Translation for LoResMT 2021

2021-08-01 · MTSummit 2021 8 · Vandan Mujadia, Dipti Misra Sharma

In this paper, we (team - oneNLP-IIITH) describe our Neural Machine Translation approaches for English-Marathi (both direction) for LoResMT-20211 . We experimented with transformer based Neural Machine Translation and ex…

Machine TranslationMORPHPOSTranslation

indic-punct: An automatic punctuation restoration and inverse text normalization framework for Indic languages

2022-03-31 · Anirudh Gupta, Neeraj Chhimwal, Ankur Dhuriya, Rishabh Gaur 외

Automatic Speech Recognition (ASR) generates text which is most of the times devoid of any punctuation. Absence of punctuation is text can affect readability. Also, down stream NLP tasks such as sentiment analysis, machi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationPunctuation Restoration+6

Evaluating the Performance of Back-translation for Low Resource English-Marathi Language Pair: CFILT-IITBombay @ LoResMT 2021

2021-08-01 · MTSummit 2021 8 · Aditya Jain, Shivam Mhaskar, Pushpak Bhattacharyya

In this paper, we discuss the details of the various Machine Translation (MT) systems that we have submitted for the English-Marathi LoResMT task. As a part of this task, we have submitted three different Neural Machine …

Machine TranslationNMTTranslation

Attentive fine-tuning of Transformers for Translation of low-resourced languages @LoResMT 2021

2021-08-19 · MTSummit 2021 8 · Karthik Puranik, Adeep Hande, Ruba Priyadharshini, Thenmozhi Durairaj 외

This paper reports the Machine Translation (MT) systems submitted by the IIITT team for the English->Marathi and English->Irish language pairs LoResMT 2021 shared task. The task focuses on getting exceptional translation…

Machine TranslationNMTTranslation

Findings of the LoResMT 2021 Shared Task on COVID and Sign Language for Low-resource Languages

2021-08-14 · MTSummit 2021 8 · Atul Kr. Ojha, Chao-Hong Liu, Katharina Kann, John Ortega 외

We present the findings of the LoResMT 2021 shared task which focuses on machine translation (MT) of COVID-19 data for both low-resource spoken and sign languages. The organization of this task was conducted as part of t…

Machine TranslationTranslation