paper-with-me

홈 › Papers

INDICXNLI: A Dataset for Studying NLI in Indic Languages

2021-11-16 · ACL ARR November 2021 11 · Anonymous

While Indic NLP has made rapid advances recently in terms of availability of corpora and pre-trained models, benchmark dataset on standard NLU tasks are limited. To this end, we introduce INDICXNLI, an NLI dataset for 11 Indic languages. It has been created by high-quality machine translation of the original English XNLI dataset and out analysis attests to the quality of INDICXNLI. By finetuning different pre-trained LMs on this INDICXNLI, we analyze various cross-lingual transfer techniques with respect to the impact of choice of language models, languages, multi-linguality, mix-language input, etc. These experiments provide us with useful insights into the behaviour of pre-trained models for a diverse set of languages. INDICXNLI will be publicly available for research.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferMachine TranslationTranslation

Similar Papers 제목 키워드 기반

IndicXNLI: Evaluating Multilingual Inference for Indian Languages

2022-04-19 · Divyanshu Aggarwal, Vivek Gupta, Anoop Kunchukuttan

While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end, we introduce IndicXNLI, an NLI dataset f…

Cross-Lingual TransferMachine TranslationTranslation

Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages

2025-08-12 · Imalsha Puranegedara, Themira Chathumina, Nisal Ranathunga, Nisansa de Silva 외 arxiv

Large Language Models (LLMs) excel in English, but their performance degrades significantly on low-resource languages (LRLs) due to English-centric training. While methods like LangBridge align LLMs with multilingual enc…

News Classification

Exploring Performance Variations in Finetuned Translators of Ultra-Low Resource Languages: Do Linguistic Differences Matter?

2025-11-27 · Isabel Gonçalves, Paulo Cavalin, Claudio Pinhanez arxiv

Finetuning pre-trained language models with small amounts of data is a commonly-used method to create translators for ultra-low resource languages such as endangered Indigenous languages. However, previous works have rep…

Prevalence and recoverability of syntactic parameters in sparse distributed memories

2015-10-21 · Jeong Joon Park, Ronnel Boettcher, Andrew Zhao, Alex Mun 외

We propose a new method, based on Sparse Distributed Memory (Kanerva Networks), for studying dependency relations between different syntactic parameters in the Principles and Parameters model of Syntax. We store data of …

Relation

Mega-COV: A Billion-Scale Dataset of 100+ Languages for COVID-19

2020-05-02 · EACL 2021 2 · Muhammad Abdul-Mageed, AbdelRahim Elmadany, El Moatez Billah Nagoudi, Dinesh Pabbi 외

We describe Mega-COV, a billion-scale dataset from Twitter for studying COVID-19. The dataset is diverse (covers 268 countries), longitudinal (goes as back as 2007), multilingual (comes in 100+ languages), and has a sign…

Misinformation