paper-with-me

Papers

How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?

2024-06-06 · Anushka Singh, Ananya B. Sai, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Mitesh M Khapra

While machine translation evaluation has been studied primarily for high-resource languages, there has been a recent interest in evaluation for low-resource languages due to the increasing availability of data and models. In this paper, we focus on a zero-shot evaluation setting focusing on low-resource Indian languages, namely Assamese, Kannada, Maithili, and Punjabi. We collect sufficient Multi-Dimensional Quality Metrics (MQM) and Direct Assessment (DA) annotations to create test sets and meta-evaluate a plethora of automatic evaluation metrics. We observe that even for learned metrics, which are known to exhibit zero-shot performance, the Kendall Tau and Pearson correlations with human annotations are only as high as 0.32 and 0.45. Synthetic data approaches show mixed results and overall do not help close the gap by much for these languages. This indicates that there is still a long way to go for low-resource evaluation.

📄 PDF Abstract BibTeX arXiv:2406.03893

Code (1)

ai4bharat/indicmt-eval 공식 구현 pytorch

Tasks

Machine Translation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Machine Translation by Projecting Text into the Same Phonetic-Orthographic Space Using a Common Encoding

2023-05-21 · Amit Kumar, Shantipriya Parida, Ajay Pratap, Anil Kumar Singh

The use of subword embedding has proved to be a major innovation in Neural Machine Translation (NMT). It helps NMT to learn better context vectors for Low Resource Languages (LRLs) so as to predict the target words by be…

Machine TranslationNMTTranslation

Zero-shot Disfluency Detection for Indian Languages

2022-10-01 · COLING 2022 10 · Rohit Kundu, Preethi Jyothi, Pushpak Bhattacharyya

Disfluencies that appear in the transcriptions from automatic speech recognition systems tend to impair the performance of downstream NLP tasks. Disfluency correction models can help alleviate this problem. However, the …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Zero-shot translation among Indian languages

2020-12-01 · loresmt (AACL) 2020 12 · Rudali Huidrom, Yves Lepage

Standard neural machine translation (NMT) allows a model to perform translation between a pair of languages. Multilingual neural machine translation (NMT), on the other hand, allows a model to perform translation between…

Machine TranslationNMTSentenceTranslation

IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS

2024-09-09 · Ashwin Sankar, Srija Anand, Praveen Srinivasa Varadhan, Sherry Thomas 외

Recent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is scarce for Indian languages due to the lack…

DenoisingSpeech Enhancementtext-to-speechText to Speech+1

IndicGEC: Powerful Models, or a Measurement Mirage?

2025-11-19 · Sowmya Vajjala arxiv

In this paper, we report the results of the TeamNRC's participation in the BHASHA-Task 1 Grammatical Error Correction shared task https://github.com/BHASHA-Workshop/IndicGEC2025/ for 5 Indian languages. Our approach, foc…

Grammatical Error Correction