MirasText: An Automatically Generated Text Corpus for Persian
Code (0)
등록된 구현이 없습니다.
Tasks
Keyword ExtractionLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
The Impact of Text Normalization on Multiword Expressions Discovery in Persian
This paper evaluates normalization procedures of Persian text for a downstream NLP task - multiword expressions (MWEs) discovery. We discuss the challenges the Persian language poses for NLP and evaluate open-source tool…
Text NormalizationPersian-WSD-Corpus: A Sense Annotated Corpus for Persian All-words Word Sense Disambiguation
Word Sense Disambiguation (WSD) is a long-standing task in Natural Language Processing(NLP) that aims to automatically identify the most relevant meaning of the words in a given context. Developing standard WSD test coll…
AllWord Sense DisambiguationPersian Wordnet Construction using Supervised Learning
This paper presents an automated supervised method for Persian wordnet construction. Using a Persian corpus and a bi-lingual dictionary, the initial links between Persian words and Princeton WordNet synsets have been gen…
ClassificationGeneral ClassificationParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice…
Text-To-Speech SynthesisSpeaker IdentificationLanguage ModellingThe First Parallel Multilingual Corpus of Persian: Toward a Persian BLARK
In this article, we have introduced the first parallel corpus of Persian with more than 10 other European languages. This article describes primary steps toward preparing a Basic Language Resources Kit (BLARK) for Persia…