Low Resource Chat Translation: A Benchmark for Hindi–English Language Pair
Chatbots are used in various sectors such as banking, healthcare, e-commerce, etc, and are mainly available in English. Machine Translation (MT) could be an effective way to develop multilingual chatbots. This paper provides a benchmark setup for Chat and QnA translation for English-Hindi, a relatively low-resource language pair. We create the English-Hindi parallel corpus consisting of both synthetic and gold standard parallel sentences from WMT20 Chat, MMD, and Flipkart QA corpora. We conduct experiments on Transfer Learning, Domain Adaptation techniques at sentence-level and context-level. We achieve BLEU scores of 57.6 and 55.5 on the En-Hi and Hi-En subsets of WMT20 Chat corpus, 51.5 and 80.4 BLEU scores on En-Hi and Hi-En subsets of MMD corpus, and 50.6 and 46.3 BLEU scores on question and answer subsets of QA corpus respectively. We also perform thorough quality tests by a very well-known e-commerce industry by deploying them to real users.
Code (1)
Tasks
Domain AdaptationMachine TranslationSentenceTransfer LearningTranslationSimilar Papers 제목 키워드 기반
Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
Multilingual LLMs support a variety of languages; however, their performance is suboptimal for low-resource languages. In this work, we emphasize the importance of continued pre-training of multilingual LLMs and the use …
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
Evaluating instruction-tuned Large Language Models (LLMs) in Hindi is challenging due to a lack of high-quality benchmarks, as direct translation of English datasets fails to capture crucial linguistic and cultural nuanc…
Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus
The linguistic diversity of India poses significant machine translation challenges, especially for underrepresented tribal languages like Bhili, which lack high-quality linguistic resources. This paper addresses the gap …
Domain GeneralizationMachine TranslationPhrase Pair Mappings for Hindi-English Statistical Machine Translation
In this paper, we present our work on the creation of lexical resources for the Machine Translation between English and Hindi. We describes the development of phrase pair mappings for our experiments and the comparative …
Machine TranslationTranslationBenchmarking of English-Hindi parallel corpora
In this paper we present several parallel corpora for English{\^a}Hindi and talk about their natures and domains. We also discuss briefly a few previous attempts in MT for translation from English to Hindi. The lack of…
BenchmarkingMachine TranslationTranslation