paper-with-me

홈 › Papers

AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages

2020-04-30 · Anoop Kunchukuttan, Divyanshu Kakwani, Satish Golla, Gokul N. C., Avik Bhattacharyya, Mitesh M. Khapra, Pratyush Kumar

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article category classification datasets for 9 languages to evaluate the embeddings. We show that the IndicNLP embeddings significantly outperform publicly available pre-trained embedding on multiple evaluation tasks. We hope that the availability of the corpus will accelerate Indic NLP research. The resources are available at https://github.com/ai4bharat-indicnlp/indicnlp_corpus.

📄 PDF Abstract BibTeX arXiv:2005.00085

Code (2)

ai4bharat-indicnlp/indicnlp_corpus 공식 구현 pytorch
csebuetnlp/banglabert pytorch

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages

2020-11-08 · Findings of the Association for Computational Linguistics 2020 · Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C. 외

In this paper, we introduce NLP resources for 11 major Indian languages from two major language families. These resources include: (a) large-scale sentence-level monolingual corpora, (b) pre-trained word embeddings, (c)…

Genre classificationMultiple-choiceMultiple Choice Question Answering (MCQA)named-entity-recognition+8

Building a Monolingual Parallel Corpus for Text Simplification Using Sentence Similarity Based on Alignment between Word Embeddings

2016-12-01 · COLING 2016 12 · Tomoyuki Kajiwara, Mamoru Komachi

Methods for text simplification using the framework of statistical machine translation have been extensively studied in recent years. However, building the monolingual parallel corpus necessary for training the model req…

Machine TranslationSentenceSentence SimilarityText Simplification+2

Estimating related words computationally using language model from the Mahabharata - an Indian epic

2023-05-09 · Vrunda Gadesha, Keyur D Joshi, Shefali Naik

'Mahabharata' is the most popular among many Indian pieces of literature referred to in many domains for completely different purposes. This text itself is having various dimension and aspects which is useful for the hum…

Language ModelingLanguage ModellingSentence

Monolingual and Parallel Corpora for Kangri Low Resource Language

2021-03-22 · Shweta Chauhan, Shefali Saxena, Philemon Daniel

In this paper we present the dataset of Himachali low resource endangered language, Kangri (ISO 639-3xnr) listed in the United Nations Educational, Scientific and Cultural Organization (UNESCO). The compilation of kangri…

Machine TranslationNMTTranslationWord Embeddings

Noisy Parallel Corpus Filtering through Projected Word Embeddings

2019-08-01 · WS 2019 8 · Murathan Kurfal{\i}, Robert {\"O}stling

We present a very simple method for parallel text cleaning of low-resource languages, based on projection of word embeddings trained on large monolingual corpora in high-resource languages. In spite of its simplicity, we…

Machine TranslationTranslationWord Embeddings