paper-with-me

홈 › Papers

Automatic Creation of a Sentence Aligned Sinhala-Tamil Parallel Corpus

2016-12-01 · WS 2016 12 · Riyafa Abdul Hameed, Nadeeshani Pathirennehelage, Anusha Ihalapathirana, Maryam Ziyad Mohamed, Surangika Ranathunga, Sanath Jayasena, Gihan Dias, Fern, S o, areka

A sentence aligned parallel corpus is an important prerequisite in statistical machine translation. However, manual creation of such a parallel corpus is time consuming, and requires experts fluent in both languages. Automatic creation of a sentence aligned parallel corpus using parallel text is the solution to this problem. In this paper, we present the first ever empirical evaluation carried out to identify the best method to automatically create a sentence aligned Sinhala-Tamil parallel corpus. Annual reports from Sri Lankan government institutions were used as the parallel text for aligning. Despite both Sinhala and Tamil being under-resourced languages, we were able to achieve an F-score value of 0.791 using a hybrid approach that makes use of a bilingual dictionary.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceTranslationWord Alignment

Similar Papers 제목 키워드 기반

Adapting the Tesseract Open-Source OCR Engine for Tamil and Sinhala Legacy Fonts and Creating a Parallel Corpus for Tamil-Sinhala-English

2021-09-13 · Charangan Vasantharajan, Laksika Tharmalingam, Uthayasanker Thayasivam

Most low-resource languages do not have the necessary resources to create even a substantial monolingual corpus. These languages may often be found in government proceedings but mainly in Portable Document Format (PDF) t…

Optical Character Recognition (OCR)

Exploiting Parallel Corpora to Improve Multilingual Embedding based Document and Sentence Alignment

2021-06-12 · Dilan Sachintha, Lakmali Piyarathna, Charith Rajitha, Surangika Ranathunga

Multilingual sentence representations pose a great advantage for low-resource languages that do not have enough data to build monolingual models on their own. These multilingual sentence representations have been separat…

Sentence

A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala

2024-12-03 · Surangika Ranathunga, Asanka Ranasinghea, Janaka Shamala, Ayodya Dandeniyaa 외

This paper presents a multi-way parallel English-Tamil-Sinhala corpus annotated with Named Entities (NEs), where Sinhala and Tamil are low-resource languages. Using pre-trained multilingual Language Models (mLMs), we est…

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine Translationnamed-entity-recognition+5

Multi-lingual Mathematical Word Problem Generation using Long Short Term Memory Networks with Enhanced Input Features

2020-05-01 · LREC 2020 5 · Vijini Liyanage, Surangika Ranathunga

A Mathematical Word Problem (MWP) differs from a general textual representation due to the fact that it is comprised of numerical quantities and units, in addition to text. Therefore, MWP generation should be carefully h…

POSTAGWord Embeddings

OasisSimp: An Open-source Asian-English Sentence Simplification Dataset

2026-03-14 · Hannah Liu, Muxin Tian, Iqra Ali, Haonan Gao 외 arxiv

Sentence simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains limited for mid-resource and low-resource …