paper-with-me

홈 › Papers

EnAsCorp1.0: English-Assamese Corpus

2020-12-01 · loresmt (AACL) 2020 12 · Sahinur Rahman Laskar, Abdullah Faiz Ur Rahman Khilji, Partha Pakray, Sivaji Bandyopadhyay

The corpus preparation is one of the important challenging task for the domain of machine translation especially in low resource language scenarios. Country like India where multiple languages exists, machine translation attempts to minimize the communication gap among people with different linguistic backgrounds. Although Google Translation covers automatic translation of various languages all over the world but it lags in some languages including Assamese. In this paper, we have developed EnAsCorp1.0, corpus of English-Assamese low resource pair where parallel and monolingual data are collected from various online sources. We have also implemented baseline systems with statistical machine translation and neural machine translation approaches for the same corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Assamese-English Bilingual Machine Translation

2014-07-08 · Kalyanee Kanchan Baruah, Pranjal Das, Abdul Hannan, Shikhar Kr. Sarma

Machine translation is the process of translating text from one language to another. In this paper, Statistical Machine Translation is done on Assamese and English language by taking their respective parallel corpus. A s…

Language ModelingLanguage ModellingMachine TranslationTranslation+1

SPRING Lab IITM's submission to Low Resource Indic Language Translation Shared Task

2024-11-01 · Hamees Sayed, Advait Joglekar, Srinivasan Umesh

We develop a robust translation model for four low-resource Indic languages: Khasi, Mizo, Manipuri, and Assamese. Our approach includes a comprehensive pipeline from data collection and preprocessing to training and eval…

Language ModellingTranslation

Development and Transcription of Assamese Speech Corpus

2013-09-27 · Himangshu Sarma, Navanath Saharia, Utpal Sharma, Smriti Kumar Sinha 외

A balanced speech corpus is the basic need for any speech processing task. In this report we describe our effort on development of Assamese speech corpus. We mainly focused on some issues and challenges faced during deve…

A Survey of Named Entity Recognition in Assamese and other Indian Languages

2014-07-09 · Gitimoni Talukdar, Pranjal Protim Borah, Arup Baruah

Named Entity Recognition is always important when dealing with major Natural Language Processing tasks such as information extraction, question-answering, machine translation, document summarization etc so in this paper …

Document SummarizationMachine Translationnamed-entity-recognitionNamed Entity Recognition+3

A Structured Approach for Building Assamese Corpus: Insights, Applications and Challenges

2012-12-01 · WS 2012 12 · Shikhar Kr. Sarma, Himadri Bharali, Ambeswar Gogoi, Ratul Deka 외