paper-with-me

Papers

A Tulu Resource for Machine Translation

2024-03-28 · Manu Narayanan, Noëmi Aepli

We present the first parallel dataset for English-Tulu translation. Tulu, classified within the South Dravidian linguistic family branch, is predominantly spoken by approximately 2.5 million individuals in southwestern India. Our dataset is constructed by integrating human translations into the multilingual machine translation resource FLORES-200. Furthermore, we use this dataset for evaluation purposes in developing our English-Tulu machine translation model. For the model's training, we leverage resources available for related South Dravidian languages. We adopt a transfer learning approach that exploits similarities between high-resource and low-resource languages. This method enables the training of a machine translation system even in the absence of parallel data between the source and target language, thereby overcoming a significant obstacle in machine translation development for low-resource languages. Our English-Tulu system, trained without using parallel English-Tulu data, outperforms Google Translate by 19 BLEU points (in September 2023). The dataset and code are available here: https://github.com/manunarayanan/Tulu-NMT.

📄 PDF Abstract BibTeX arXiv:2403.19142

Code (1)

manunarayanan/tulu-nmt 공식 구현

Tasks

Machine TranslationNMTTransfer LearningTranslation

Similar Papers 제목 키워드 기반

TULUN: Transparent and Adaptable Low-resource Machine Translation

2025-05-24 · Raphaël Merx, Hanna Suominen, Lois Hong, Nick Thieberger 외

Machine translation (MT) systems that support low-resource languages often struggle on specialized domains. While researchers have proposed various techniques for domain adaptation, these approaches typically require mod…

Domain AdaptationLanguage ModelingLanguage ModellingLarge Language Model+2

Corpus Creation for Sentiment Analysis in Code-Mixed Tulu Text

2022-06-01 · SIGUL (LREC) 2022 6 · Asha Hegde, Mudoor Devadas Anusha, Sharal Coelho, Hosahalli Lakshmaiah Shashirekha 외

Sentiment Analysis (SA) employing code-mixed data from social media helps in getting insights to the data and decision making for various applications. One such application is to analyze users’ emotions from comments of …

Decision MakingSentiment Analysis

Overview of the Shared Task on Machine Translation in Dravidian Languages

2022-05-01 · DravidianLangTech (ACL) 2022 5 · Anand Kumar Madasamy, Asha Hegde, Shubhanker Banerjee, Bharathi Raja Chakravarthi 외

This paper presents an outline of the shared task on translation of under-resourced Dravidian languages at DravidianLangTech-2022 workshop to be held jointly with ACL 2022. A description of the datasets used, approach ta…

Machine TranslationTranslation

Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification

2025-08-15 · Anusha M D, Deepthi Vikram, Bharathi Raja Chakravarthi, Parameshwar R Hegde arxiv

Tulu, a low-resource Dravidian language predominantly spoken in southern India, has limited computational resources despite its growing digital presence. This study presents the first benchmark dataset for Offensive Lang…

Language Identification

BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages

2024-12-05 · Vandan Mujadia, Dipti Misra Sharma

This paper focuses on developing translation models and related applications for 36 Indian languages, including Assamese, Awadhi, Bengali, Bhojpuri, Braj, Bodo, Dogri, English, Konkani, Gondi, Gujarati, Hindi, Hinglish, …

Automatic Post-EditingData AugmentationMachine TranslationTranslation