paper-with-me

Papers

NLIP_Lab-IITH Multilingual MT System for WAT24 MT Shared Task

2024-10-17 · Maharaj Brahma, Pramit Sahoo, Maunendra Sankar Desarkar

This paper describes NLIP Lab's multilingual machine translation system for the WAT24 shared task on multilingual Indic MT task for 22 scheduled languages belonging to 4 language families. We explore pre-training for Indic languages using alignment agreement objectives. We utilize bi-lingual dictionaries to substitute words from source sentences. Furthermore, we fine-tuned language direction-specific multilingual translation models using small and high-quality seed data. Our primary submission is a 243M parameters multilingual translation model covering 22 Indic languages. In the IN22-Gen benchmark, we achieved an average chrF++ score of 46.80 and 18.19 BLEU score for the En-Indic direction. In the Indic-En direction, we achieved an average chrF++ score of 56.34 and 30.82 BLEU score. In the In22-Conv benchmark, we achieved an average chrF++ score of 43.43 and BLEU score of 16.58 in the En-Indic direction, and in the Indic-En direction, we achieved an average of 52.44 and 29.77 for chrF++ and BLEU respectively. Our model\footnote{Our code and models are available at \url{https://github.com/maharajbrahma/WAT2024-MultiIndicMT}} is competitive with IndicTransv1 (474M parameter model).

📄 PDF Abstract BibTeX arXiv:2410.13443

Code (1)

maharajbrahma/wat2024-multiindicmt 공식 구현

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

NLIP_Lab-IITH Low-Resource MT System for WMT24 Indic MT Shared Task

2024-10-04 · Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar

In this paper, we describe our system for the WMT 24 shared task of Low-Resource Indic Language Translation. We consider eng $\leftrightarrow$ {as, kha, lus, mni} as participating language pairs. In this shared task, we …

IIT(BHU)--IIITH at CoNLL--SIGMORPHON 2018 Shared Task on Universal Morphological Reinflection

2018-10-01 · CONLL 2018 10 · Abhishek Sharma, Ganesh Katrapati, Dipti Misra Sharma
Feature EngineeringMorphological Inflection

PreCogIIITH at HinglishEval : Leveraging Code-Mixing Metrics & Language Model Embeddings To Estimate Code-Mix Quality

2022-06-16 · Prashant Kodali, Tanmay Sachan, Akshay Goindani, Anmol Goel 외

Code-Mixing is a phenomenon of mixing two or more languages in a speech event and is prevalent in multilingual societies. Given the low-resource nature of Code-Mixing, machine generation of code-mixed text is a prevalent…

Data AugmentationLanguage ModelingLanguage Modelling

IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation

2025-06-05 · Bhavana Akkiraju, Aishwarya Pothula, Santosh Kesiraju, Anil Kumar Vuppala

This paper presents the submission of IIITH-BUT to the IWSLT 2025 shared task on speech translation for the low-resource Bhojpuri-Hindi language pair. We explored the impact of hyperparameter optimisation and data augmen…

Data AugmentationTranslation

Precog-LTRC-IIITH at GermEval 2021: Ensembling Pre-Trained Language Models with Feature Engineering

2021-09-01 · GermEval 2021 9 · T. H. Arjun, Arvindh A., Kumaraguru Ponnurangam

We describe our participation in all the subtasks of the Germeval 2021 shared task on the identification of Toxic, Engaging, and Fact-Claiming Comments. Our system is an ensemble of state-of-the-art pre-trained models fi…

Data AugmentationFeature Engineering