paper-with-me

홈 › Papers

Enhanced Urdu Word Segmentation using Conditional Random Fields and Morphological Context Features

2020-07-01 · WS 2020 7 · Aamir Farhan, Mashrukh Islam, Dipti Misra Sharma

Word segmentation is a fundamental task for most of the NLP applications. Urdu adopts Nastalique writing style which does not have a concept of space. Furthermore, the inherent non-joining attributes of certain characters in Urdu create spaces within a word while writing in digital format. Thus, Urdu not only has space omission but also space insertion issues which make the word segmentation task challenging. In this paper, we improve upon the results of Zia, Raza and Athar (2018) by using a manually annotated corpus of 19,651 sentences along with morphological context features. Using the Conditional Random Field sequence modeler, our model achieves F 1 score of 0.98 for word boundary identification and 0.92 for sub-word boundary identification tasks. The results demonstrated in this paper outperform the state-of-the-art methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Segmentation

Similar Papers 제목 키워드 기반

Urdu Word Segmentation using Conditional Random Fields (CRFs)

2018-06-14 · COLING 2018 8 · Haris Bin Zia, Agha Ali Raza, Awais Athar

State-of-the-art Natural Language Processing algorithms rely heavily on efficient word segmentation. Urdu is amongst languages for which word segmentation is a complex task as it exhibits space omission as well as space …

Segmentation

MUTEX: Leveraging Multilingual Transformers and Conditional Random Fields for Enhanced Urdu Toxic Span Detection

2026-03-05 · Inayat Arshad, Fajar Saleem, Ijaz Hussain arxiv

Urdu toxic span detection remains limited because most existing systems rely on sentence-level classification and fail to identify the specific toxic spans within those text. It is further exacerbated by the multiple fac…

Context based Roman-Urdu to Urdu Script Transliteration System

2021-09-29 · H Muhammad Shakeel, Rashid Khan, Muhammad Waheed

Now a day computer is necessary for human being and it is very useful in many fields like search engine, text processing, short messaging services, voice chatting and text recognition. Since last many years there are man…

Transliteration

Co-occurrences using Fasttext embeddings for word similarity tasks in Urdu

2021-02-22 · Usama Khalid, Aizaz Hussain, Muhammad Umair Arshad, Waseem Shahzad 외

Urdu is a widely spoken language in South Asia. Though immoderate literature exists for the Urdu language still the data isn't enough to naturally process the language by NLP techniques. Very efficient language models ex…

Word EmbeddingsWord Similarity

Khmer Word Segmentation Using Conditional Random Fields

2015-10-15 · Vichet Chea, Ye Kyaw Thu, Chenchen Ding, Masao Utiyama 외

Word Segmentation is a critical task that is the foundation of much natural language processing research. This paper is a study of Khmer word segmentation using an approach based on conditional random fields (CRFs). A…

SegmentationText SegmentationTranslation