paper-with-me

Papers

UDAAN: Machine Learning based Post-Editing tool for Document Translation

2022-03-03 · Ayush Maheshwari, Ajay Ravindran, Venkatapathy Subramanian, Ganesh Ramakrishnan

We introduce UDAAN, an open-source post-editing tool that can reduce manual editing efforts to quickly produce publishable-standard documents in several Indic languages. UDAAN has an end-to-end Machine Translation (MT) plus post-editing pipeline wherein users can upload a document to obtain raw MT output. Further, users can edit the raw translations using our tool. UDAAN offers several advantages: a) Domain-aware, vocabulary-based lexical constrained MT. b) source-target and target-target lexicon suggestions for users. Replacements are based on the source and target texts lexicon alignment. c) Translation suggestions are based on logs created during user interaction. d) Source-target sentence alignment visualisation that reduces the cognitive load of users during editing. e) Translated outputs from our tool are available in multiple formats: docs, latex, and PDF. We also provide the facility to use around 100 in-domain dictionaries for lexicon-aware machine translation. Although we limit our experiments to English-to-Hindi translation, our tool is independent of the source and target languages. Experimental results based on the usage of the tools and users feedback show that our tool speeds up the translation time by approximately a factor of three compared to the baseline method of translating documents from scratch. Our tool is available for both Windows and Linux platforms. The tool is open-source under MIT license, and the source code can be accessed from our website at https://www.udaanproject.org. Demonstration and tutorial videos for various features of our tool can be accessed at https://www.youtube.com/channel/UClfK7iC8J7b22bj3GwAUaCw. Our MT pipeline can be accessed at https://udaaniitb.aicte-india.org/udaan/translate/.

📄 PDF Abstract BibTeX arXiv:2203.01644

Code (1)

IITB-OpenOCRCorrect/iitb-openocr-digit-tool 공식 구현

Tasks

BIG-bench Machine LearningDocument TranslationMachine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

A Tool for Facilitating OCR Postediting in Historical Documents

2020-04-23 · LREC 2020 5 · Alberto Poncelas, Mohammad Aboomar, Jan Buts, James Hadley 외

Optical character recognition (OCR) for historical documents is a complex procedure subject to a unique set of material issues, including inconsistencies in typefaces and low quality scanning. Consequently, even the most…

Language ModelingLanguage ModellingOptical Character RecognitionOptical Character Recognition (OCR)

PET: a Tool for Post-editing and Assessing Machine Translation

2012-05-01 · LREC 2012 5 · Wilker Aziz, Sheila Castilho, Lucia Specia

Given the significant improvements in Machine Translation (MT) quality and the increasing demand for translations, post-editing of automatic translations is becoming a popular practice in the translation industry. It has…

Machine TranslationSentenceTranslation

Translator2Vec: Understanding and Representing Human Post-Editors

2019-07-24 · António Góis, André F. T. Martins

The combination of machines and humans for translation is effective, with many studies showing productivity gains when humans post-edit machine-translated output instead of translating from scratch. To take full advantag…

Translation

IntelliCAT: Intelligent Machine Translation Post-Editing with Quality Estimation and Translation Suggestion

2021-05-25 · ACL 2021 5 · Dongjun Lee, Junhyeong Ahn, Heesoo Park, Jaemin Jo

We present IntelliCAT, an interactive translation interface with neural models that streamline the post-editing process on machine translation output. We leverage two quality estimation (QE) models at different granulari…

Machine TranslationSentenceTranslation

Exploring the Importance of Source Text in Automatic Post-Editing for Context-Aware Machine Translation

2021-05-01 · NoDaLiDa 2021 5 · Chaojun Wang, Christian Hardmeier, Rico Sennrich

Accurate translation requires document-level information, which is ignored by sentence-level machine translation. Recent work has demonstrated that document-level consistency can be improved with automatic post-editing (…

Automatic Post-EditingMachine TranslationSentenceTranslation