First Attempt at Building Parallel Corpora for Machine Translation of Northeast India's Very Low-Resource Languages
This paper presents the creation of initial bilingual corpora for thirteen very low-resource languages of India, all from Northeast India. It also presents the results of initial translation efforts in these languages. It creates the first-ever parallel corpora for these languages and provides initial benchmark neural machine translation results for these languages. We intend to extend these corpora to include a large number of low-resource Indian languages and integrate the effort with our prior work with African and American-Indian languages to create corpora covering a large number of languages from across the world.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Leveraging Multilingual News Websites for Building a Kurdish Parallel Corpus
Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems thanks to deep learning methods, p…
ArticlesMachine TranslationTranslationTransliterationExploring Word Alignment towards an Efficient Sentence Aligner for Filipino and Cebuano Languages
Building a robust machine translation (MT) system requires a large amount of parallel corpus which is an expensive resource for low-resourced languages. The two major languages being spoken in the Philippines which are F…
Machine TranslationSentenceTranslationWord AlignmentBuilding Machine Translation System for Software Product Descriptions Using Domain-specific Sub-corpora Extraction
Building Machine Translation systems for a specific domain requires a sufficiently large and good quality parallel corpus in that domain. However, this is a bit challenging task due to the lack of parallel data in many d…
Machine TranslationSentenceSentence EmbeddingSentence-Embedding+1Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…
General ClassificationMachine TranslationRetrievalTranslationBuilding Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such…
ArticlesMachine TranslationRetrievalSentence+1