paper-with-me

Papers

Urdu Summary Corpus

2016-05-01 · LREC 2016 5

Language resources, such as corpora, are important for various natural language processing tasks. Urdu has millions of speakers around the world but it is under-resourced in terms of standard evaluation resources. This paper reports the construction of a benchmark corpus for Urdu summaries (abstracts) to facilitate the development and evaluation of single document summarization systems for Urdu language. In Urdu, space does not always mark word boundary. Therefore, we created two versions of the same corpus. In the first version, words are separated by space. In contrast, proper word boundaries are manually tagged in the second version. We further apply normalization, part-of-speech tagging, morphological analysis, lemmatization, and stemming for the articles and their summaries in both versions. In order to apply these annotations, we re-implemented some NLP tools for Urdu. We provide Urdu Summary Corpus, all these annotations and the needed software tools (as open-source) for researchers to run experiments and to evaluate their work including but not limited to single-document summarization task.

📄 PDF Abstract BibTeX

Code (1)

humsha/USCorpus 공식 구현

Tasks

ArticlesDocument SummarizationLemmatizationMorphological AnalysisPart-Of-Speech Tagging

Similar Papers 제목 키워드 기반

Abstractive Summary Generation for the Urdu Language

2023-05-25 · Ali Raza, Hadia Sultan Raja, Usman Maratib

Abstractive summary generation is a challenging task that requires the model to comprehend the source text and generate a concise and coherent summary that captures the essential information. In this paper, we explore th…

Decoder

Learning Trilingual Dictionaries for Urdu -- Roman Urdu -- English

2019-08-01 · WS 2019 8 · Moiz Rauf, Sebastian Pad{\'o}

In this paper, we present an effort to generate a joint Urdu, Roman Urdu and English trilingual lexicon using automated methods. We make a case for using statistical machine translation approaches and parallel corpora fo…

Machine TranslationTranslationWord Alignment

Clustering Urdu News Using Headlines

2015-09-27 · 23 2015 9 · Samia Khaliq, Waheed Iqbal, Faisal Bukhari, Kamran Malik

This paper that proposes and evaluates a new algorithm to automatically cluster Urdu news from different news agencies. The task is challenging because there are no language processing libraries for the Urdu language. Th…

ClusteringInformation RetrievalText Clustering

A Tagged Corpus and a Tagger for Urdu

2014-05-01 · LREC 2014 5 · Bushra Jawaid, Amir Kamran, Ond{\v{r}}ej Bojar

In this paper, we describe a release of a sizeable monolingual Urdu corpus automatically tagged with part-of-speech tags. We extend the work of Jawaid and Bojar (2012) who use three different taggers and then apply a vot…

Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features

2025-10-28 · Unzela Talpur, Zafi Sherhan Syed, Muhammad Shehram Shah Syed, Abbas Shah Syed arxiv

Speech Emotion Recognition (SER) is a key affective computing technology that enables emotionally intelligent artificial intelligence. While SER is challenging in general, it is particularly difficult for low-resource la…

Speech Emotion Recognition