paper-with-me

홈 › Papers

Universal Dependencies for Persian

2016-05-01 · LREC 2016 5 · Mojgan Seraji, Filip Ginter, Joakim Nivre

The Persian Universal Dependency Treebank (Persian UD) is a recent effort of treebanking Persian with Universal Dependencies (UD), an ongoing project that designs unified and cross-linguistically valid grammatical representations including part-of-speech tags, morphological features, and dependency relations. The Persian UD is the converted version of the Uppsala Persian Dependency Treebank (UPDT) to the universal dependencies framework and consists of nearly 6,000 sentences and 152,871 word tokens with an average sentence length of 25 words. In addition to the universal dependencies syntactic annotation guidelines, the two treebanks differ in tokenization. All words containing unsegmented clitics (pronominal and copula clitics) annotated with complex labels in the UPDT have been separated from the clitics and appear with distinct labels in the Persian UD. The treebank has its original syntactic annotation scheme based on Stanford Typed Dependencies. In this paper, we present the approaches taken in the development of the Persian UD.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentencevalid

Similar Papers 제목 키워드 기반

The Persian Dependency Treebank Made Universal

2020-09-21 · LREC 2022 6 · Mohammad Sadegh Rasooli, Pegah Safari, Amirsaeid Moloodi, Alireza Nourian

We describe an automatic method for converting the Persian Dependency Treebank (Rasooli et al, 2013) to Universal Dependencies. This treebank contains 29107 sentences. Our experiments along with manual linguistic analysi…

Informal Persian Universal Dependency Treebank

2022-01-10 · LREC 2022 6 · Roya Kabiri, Simin Karimi, Mihai Surdeanu

This paper presents the phonological, morphological, and syntactic distinctions between formal and informal Persian, showing that these two variants have fundamental differences that cannot be attributed solely to pronun…

Persian Abstract Meaning Representation: Annotation Guidelines and Gold Standard Dataset

2022-05-16 · Reza Takhshid, Tara Azin, Razieh Shojaei, Mohammad bahrani

This paper introduces the Persian Abstract Meaning Representation (AMR) guidelines, a detailed guide for annotating Persian sentences with AMR, focusing on the necessary adaptations to fit Persian's unique syntactic stru…

Abstract Meaning RepresentationSentenceTranslation

A Persian Treebank with Stanford Typed Dependencies

2014-05-01 · LREC 2014 5 · Mojgan Seraji, Carina Jahani, Be{\'a}ta Megyesi, Joakim Nivre

We present the Uppsala Persian Dependency Treebank (UPDT) with a syntactic annotation scheme based on Stanford Typed Dependencies. The treebank consists of 6,000 sentences and 151,671 tokens with an average sentence leng…

ArticlesCultural Vocal Bursts Intensity PredictionSentence

A Computational Approach to Language Contact -- A Case Study of Persian

2026-01-28 · Ali Basirat, Danial Namazifard, Navid Baradaran Hemmati arxiv

We investigate structural traces of language contact in the intermediate representations of a monolingual language model. Focusing on Persian (Farsi) as a historically contact-rich language, we probe the representations …