paper-with-me

Papers

First Attempt at Building Parallel Corpora for Machine Translation of Northeast India's Very Low-Resource Languages

2023-12-08 · Atnafu Lambebo Tonja, Melkamu Mersha, Ananya Kalita, Olga Kolesnikova, Jugal Kalita

This paper presents the creation of initial bilingual corpora for thirteen very low-resource languages of India, all from Northeast India. It also presents the results of initial translation efforts in these languages. It creates the first-ever parallel corpora for these languages and provides initial benchmark neural machine translation results for these languages. We intend to extend these corpora to include a large number of low-resource Indian languages and integrate the effort with our prior work with African and American-Indian languages to create corpora covering a large number of languages from across the world.

📄 PDF Abstract BibTeX arXiv:2312.04764

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Leveraging Multilingual News Websites for Building a Kurdish Parallel Corpus

2020-10-04 · Sina Ahmadi, Hossein Hassani, Daban Q. Jaff

Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems thanks to deep learning methods, p…

ArticlesMachine TranslationTranslationTransliteration

Exploring Word Alignment towards an Efficient Sentence Aligner for Filipino and Cebuano Languages

2022-10-01 · loresmt (COLING) 2022 10 · Jenn Leana Fernandez, Kristine Mae M. Adlaon

Building a robust machine translation (MT) system requires a large amount of parallel corpus which is an expensive resource for low-resourced languages. The two major languages being spoken in the Philippines which are F…

Machine TranslationSentenceTranslationWord Alignment

Building Machine Translation System for Software Product Descriptions Using Domain-specific Sub-corpora Extraction

2022-09-01 · AMTA 2022 9 · Pintu Lohar, Sinead Madden, Edmond O’Connor, Maja Popovic 외

Building Machine Translation systems for a specific domain requires a sufficiently large and good quality parallel corpus in that domain. However, this is a bit challenging task due to the lack of parallel data in many d…

Machine TranslationSentenceSentence EmbeddingSentence-Embedding+1

Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics

2015-11-18 · Krzysztof Wołk, Emilia Rejmund, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…

General ClassificationMachine TranslationRetrievalTranslation

Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs

2015-09-29 · Krzysztof Wołk, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such…

ArticlesMachine TranslationRetrievalSentence+1