paper-with-me

홈 › Papers

SiPOS: A Benchmark Dataset for Sindhi Part-of-Speech Tagging

2021-09-01 · RANLP 2021 9 · Wazir Ali, Zenglin Xu, Jay Kumar

In this paper, we introduce the SiPOS dataset for part-of-speech tagging in the low-resource Sindhi language with quality baselines. The dataset consists of more than 293K tokens annotated with sixteen universal part-of-speech categories. Two experienced native annotators annotated the SiPOS using the Doccano text annotation tool with an inter-annotation agreement of 0.872. We exploit the conditional random field, the popular bidirectional long-short-term memory neural model, and self-attention mechanism with various settings to evaluate the proposed dataset. Besides pre-trained GloVe and fastText representation, the character-level representations are incorporated to extract character-level information using the bidirectional long-short-term memory encoder. The high accuracy of 96.25% is achieved with the task-specific joint word-level and character-level representations. The SiPOS dataset is likely to be a significant resource for the low-resource Sindhi language.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech Taggingtext annotation

Similar Papers 제목 키워드 기반

Part of Speech Tagging for a Resource Poor Language : Sindhi in Devanagari Script using HMM and CRF

2021-12-01 · ICON 2021 12 · Bharti Nathani, Nisheeth Joshi

Part of speech tagging is a pre-processing step of various NLP applications. Mainly it is used in Machine Translation. This research proposes two POS taggers, i.e., an HMM-based and CRF based tagger. To develop this tagg…

Machine TranslationPart-Of-Speech TaggingPOSTranslation

A Finite-State Morphological Analyser for Sindhi

2016-05-01 · LREC 2016 5 · Raveesh Motlani, Francis Tyers, Dipti Sharma

Morphological analysis is a fundamental task in natural-language processing, which is used in other NLP applications such as part-of-speech tagging, syntactic parsing, information retrieval, machine translation, etc. In …

Information RetrievalLEMMAMachine TranslationMorphological Analysis+3

SiNER: A Large Dataset for Sindhi Named Entity Recognition

2020-05-01 · LREC 2020 5 · Wazir Ali, Junyu Lu, Zenglin Xu

We introduce the SiNER: a named entity recognition (NER) dataset for low-resourced Sindhi language with quality baselines. It contains 1,338 news articles and more than 1.35 million tokens collected from Kawish and Awami…

Articlesnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Phonetic based SoundEx & ShapeEx algorithm for Sindhi Spell Checker System

2014-05-13 · Zeeshan Bhatti, Ahmad Waqas, Imdad Ali Ismaili, Dil Nawaz Hakro 외

This paper presents a novel combinational phonetic algorithm for Sindhi Language, to be used in developing Sindhi Spell Checker which has yet not been developed prior to this work. The compound textual forms and glyphs o…

Design & Development of the Graphical User Interface for Sindhi Language

2014-01-07 · Imdad Ali Ismaili, Zeeshan Bhatti, Azhar Ali Shah

This paper describes the design and implementation of a Unicode-based GUISL (Graphical User Interface for Sindhi Language). The idea is to provide a software platform to the people of Sindh as well as Sindhi diasporas li…