Automatic Creation of a Sentence Aligned Sinhala-Tamil Parallel Corpus
A sentence aligned parallel corpus is an important prerequisite in statistical machine translation. However, manual creation of such a parallel corpus is time consuming, and requires experts fluent in both languages. Automatic creation of a sentence aligned parallel corpus using parallel text is the solution to this problem. In this paper, we present the first ever empirical evaluation carried out to identify the best method to automatically create a sentence aligned Sinhala-Tamil parallel corpus. Annual reports from Sri Lankan government institutions were used as the parallel text for aligning. Despite both Sinhala and Tamil being under-resourced languages, we were able to achieve an F-score value of 0.791 using a hybrid approach that makes use of a bilingual dictionary.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSentenceTranslationWord AlignmentSimilar Papers 제목 키워드 기반
Adapting the Tesseract Open-Source OCR Engine for Tamil and Sinhala Legacy Fonts and Creating a Parallel Corpus for Tamil-Sinhala-English
Most low-resource languages do not have the necessary resources to create even a substantial monolingual corpus. These languages may often be found in government proceedings but mainly in Portable Document Format (PDF) t…
Optical Character Recognition (OCR)Exploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence Alignment
Multilingual sentence representations pose a great advantage for low-resource languages that do not have enough data to build monolingual models on their own. These multilingual sentence representations have been separat…
SentenceA Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala
This paper presents a multi-way parallel English-Tamil-Sinhala corpus annotated with Named Entities (NEs), where Sinhala and Tamil are low-resource languages. Using pre-trained multilingual Language Models (mLMs), we est…
Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine Translationnamed-entity-recognition+5Multi-lingual Mathematical Word Problem Generation using Long Short Term Memory Networks with Enhanced Input Features
A Mathematical Word Problem (MWP) differs from a general textual representation due to the fact that it is comprised of numerical quantities and units, in addition to text. Therefore, MWP generation should be carefully h…
POSTAGWord EmbeddingsOasisSimp: An Open-source Asian-English Sentence Simplification Dataset
Sentence simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains limited for mid-resource and low-resource …