paper-with-me

Papers

A Basic Language Resource Kit for Persian

2012-05-01 · LREC 2012 5 · Mojgan Seraji, Be{\'a}ta Megyesi, Joakim Nivre

Persian with its about 100,000,000 speakers in the world belongs to the group of languages with less developed linguistically annotated resources and tools. The few existing resources and tools are neither open source nor freely available. Thus, our goal is to develop open source resources such as corpora and treebanks, and tools for data-driven linguistic analysis of Persian. We do this by exploring the reusability of existing resources and adapting state-of-the-art methods for the linguistic annotation. We present fully functional tools for text normalization, sentence segmentation, tokenization, part-of-speech tagging, and parsing. As for resources, we describe the Uppsala PErsian Corpus (UPEC) which is a modified version of the Bijankhan corpus with additional sentence segmentation and consistent tokenization modified for more appropriate syntactic annotation. The corpus consists of 2,782,109 tokens and is annotated with parts of speech and morphological features. A treebank is derived from UPEC with an annotation scheme based on Stanford Typed Dependencies and is planned to consist of 10,000 sentences of which 215 have already been annotated. Keywords: BLARK for Persian, PoS tagged corpus, Persian treebank

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech TaggingPOSSentenceSentence segmentationText Normalization

Similar Papers 제목 키워드 기반

The First Parallel Multilingual Corpus of Persian: Toward a Persian BLARK

2014-04-17 · Behrang Qasemizadeh, Saeed Rahimi, Behrooz Mahmoodi Bakhtiari

In this article, we have introduced the first parallel corpus of Persian with more than 10 other European languages. This article describes primary steps toward preparing a Basic Language Resources Kit (BLARK) for Persia…

PersianLLaMA: Towards Building First Persian Large Language Model

2023-12-25 · Mohammad Amin Abbasi, Arash Ghafouri, Mahdi Firouzmandi, Hassan Naderi 외

Despite the widespread use of the Persian language by millions globally, limited efforts have been made in natural language processing for this language. The use of large language models as effective tools in various nat…

Language ModelingLanguage ModellingLarge Language ModelMachine Translation+5

Winning with Less for Low Resource Languages: Advantage of Cross-Lingual English_Persian Argument Mining Model over LLM Augmentation

2025-11-25 · Ali Jahan, Masood Ghayoomi, Annette Hautli-Janisz arxiv

Argument mining is a subfield of natural language processing to identify and extract the argument components, like premises and conclusions, within a text and to recognize the relations between them. It reveals the logic…

Argument Mining

Tajik-Farsi Persian Transliteration Using Statistical Machine Translation

2012-05-01 · LREC 2012 5 · Chris Irwin Davis

Tajik Persian is a dialect of Persian spoken primarily in Tajikistan and written with a modified Cyrillic alphabet. Iranian Persian, or Farsi, as it is natively called, is the lingua franca of Iran and is written with th…

Machine TranslationTranslationTransliteration

Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning

2025-12-08 · Amir Mohammad Akhlaghi, Amirhossein Shabani, Mostafa Abdolmaleki, Saeed Reza Kheradpisheh arxiv

The democratization of AI is currently hindered by the immense computational costs required to train Large Language Models (LLMs) for low-resource languages. This paper presents Persian-Phi, a 3.8B parameter model that c…

parameter-efficient fine-tuningContinual Pretraining