paper-with-me

Papers

A Proposition Bank of Urdu

2016-05-01 · LREC 2016 5 · Maaz Anwar, Riyaz Ahmad Bhat, Dipti Sharma, Ashwini Vaidya, Martha Palmer, Tafseer Ahmed Khan

This paper describes our efforts for the development of a Proposition Bank for Urdu, an Indo-Aryan language. Our primary goal is the labeling of syntactic nodes in the existing Urdu dependency Treebank with specific argument labels. In essence, it involves annotation of predicate argument structures of both simple and complex predicates in the Treebank corpus. We describe the overall process of building the PropBank of Urdu. We discuss various statistics pertaining to the Urdu PropBank and the issues which the annotators encountered while developing the PropBank. We also discuss how these challenges were addressed to successfully expand the PropBank corpus. While reporting the Inter-annotator agreement between the two annotators, we show that the annotators share similar understanding of the annotation guidelines and of the linguistic phenomena present in the language. The present size of this Propbank is around 180,000 tokens which is double-propbanked by the two annotators for simple predicates. Another 100,000 tokens have been annotated for complex predicates of Urdu.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dependency Parsing for Urdu: Resources, Conversions and Learning

2020-05-01 · LREC 2020 5 · Toqeer Ehsan, Miriam Butt

This paper adds to the available resources for the under-resourced language Urdu by converting different types of existing treebanks for Urdu into a common format that is based on Universal Dependencies. We present compa…

Dependency ParsingWord Embeddings

The CLE Urdu POS Tagset

2014-05-01 · LREC 2014 5 · Saba Urooj, Sarmad Hussain, Asad Mustafa, Rahila Parveen 외

The paper presents a design schema and details of a new Urdu POS tagset. This tagset is designed due to challenges encountered in working with existing tagsets for Urdu. It uses tags that judiciously incorporate informat…

Machine TranslationPOSTAG

Adapting Predicate Frames for Urdu PropBanking

2014-10-01 · WS 2014 10 · Riyaz Ahmad Bhat, Naman Jain, Ashwini Vaidya, Martha Palmer 외

Dependency Treebank of Urdu and its Evaluation

2012-07-01 · WS 2012 7 · Riyaz Ahmad Bhat, Dipti Misra Sharma

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

2026-06-05 · Ahmer Tabassum, Sarfraz Ahmad, Hasan Iqbal, Owais Aijaz 외 arxiv

Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark built from native educational sources. …