paper-with-me

Papers

HmBlogs: A big general Persian corpus

2021-11-03 · Hamzeh Motahari Khansari, Mehrnoush Shamsfard

This paper introduces the hmBlogs corpus for Persian, as a low resource language. This corpus has been prepared based on a collection of nearly 20 million blog posts over a period of about 15 years from a space of Persian blogs and includes more than 6.8 billion tokens. It can be claimed that this corpus is currently the largest Persian corpus that has been prepared independently for the Persian language. This corpus is presented in both raw and preprocessed forms, and based on the preprocessed corpus some word embedding models are produced. By the provided models, the hmBlogs is compared with some of the most important corpora available in Persian, and the results show the superiority of the hmBlogs corpus over the others. These evaluations also present the importance and effects of corpora, evaluation datasets, model production methods, different hyperparameters and even the evaluation methods. In addition to evaluating the corpus and its produced language models, this research also presents a semantic analogy dataset.

📄 PDF Abstract BibTeX arXiv:2111.02362

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FaBERT: Pre-training BERT on Persian Blogs

2024-02-09 · Mostafa Masumi, Seyed Soroush Majd, Mehrnoush Shamsfard, Hamid Beigy

We introduce FaBERT, a Persian BERT-base model pre-trained on the HmBlogs corpus, encompassing both informal and formal Persian texts. FaBERT is designed to excel in traditional Natural Language Understanding (NLU) tasks…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language Inference+5

Persian-WSD-Corpus: A Sense Annotated Corpus for Persian All-words Word Sense Disambiguation

2021-07-04 · Hossein Rouhizadeh, Mehrnoush Shamsfard, Vahideh Tajalli, Masoud Rouhziadeh

Word Sense Disambiguation (WSD) is a long-standing task in Natural Language Processing(NLP) that aims to automatically identify the most relevant meaning of the words in a given context. Developing standard WSD test coll…

AllWord Sense Disambiguation

The First Parallel Multilingual Corpus of Persian: Toward a Persian BLARK

2014-04-17 · Behrang Qasemizadeh, Saeed Rahimi, Behrooz Mahmoodi Bakhtiari

In this article, we have introduced the first parallel corpus of Persian with more than 10 other European languages. This article describes primary steps toward preparing a Basic Language Resources Kit (BLARK) for Persia…

The Persian Piano Corpus: A Collection Of Instrument-Based Feature Extracted Data Considering Dastgah

2023-11-18 · Parsa Rasouli, Azam Bastanfard

The research in the field of music is rapidly growing, and this trend emphasizes the need for comprehensive data. Though researchers have made an effort to contribute their own datasets, many data collections lack the re…

Persian SemCor: A Bag of Word Sense Annotated Corpus for the Persian Language

2021-01-01 · EACL (GWC) 2021 1 · Hossein Rouhizadeh, Mehrnoush Shamsfard, Mahdi Dehghan, Masoud Rouhizadeh

Supervised approaches usually achieve the best performance in the Word Sense Disambiguation problem. However, the unavailability of large sense annotated corpora for many low-resource languages make these approaches inap…

Word Sense Disambiguation