Unknown Words Analysis in POS tagging of Sinhala Language
Part of Speech (POS) is a very vital topic in Natural Language Processing (NLP) task in any language, which involves analysing the construction of the language, behaviours and the dynamics of the language, the knowledge that could be utilized in computational linguistics analysis and automation applications. In this context, dealing with unknown words (words do not appear in the lexicon referred as unknown words) is also an important task, since growing NLP systems are used in more and more new applications. One aid of predicting lexical categories of unknown words is the use of syntactical knowledge of the language. The distinction between open class words and closed class words together with syntactical features of the language used in this research to predict lexical categories of unknown words in the tagging process. An experiment is performed to investigate the ability of the approach to parse unknown words using syntactical knowledge without human intervention. This experiment shows that the performance of the tagging process is enhanced when word class distinction is used together with syntactic rules to parse sentences containing unknown words in Sinhala language.
Code (0)
등록된 구현이 없습니다.
Tasks
POSPOS TaggingSimilar Papers 제목 키워드 기반
Hidden Markov Model Based Part of Speech Tagger for Sinhala Language
In this paper we present a fundamental lexical semantics of Sinhala language and a Hidden Markov Model (HMM) based Part of Speech (POS) Tagger for Sinhala language. In any Natural Language processing task, Part of Speech…
POSTAGComprehensive Part-Of-Speech Tag Set and SVM based POS Tagger for Sinhala
This paper presents a new comprehensive multi-level Part-Of-Speech tag set and a Support Vector Machine based Part-Of-Speech tagger for the Sinhala language. The currently available tag set for Sinhala has two limitation…
POSTAGSinhala Language Corpora and Stopwords from a Decade of Sri Lankan Facebook
This paper presents two colloquial Sinhala language corpora from the language efforts of the Data, Analysis and Policy team of LIRNEasia, as well as a list of algorithmically derived stopwords. The larger of the two corp…
Building a Linguistic Resource : A Word Frequency List for Sinhala
A word frequency list is a list of unique words in a language along with their frequency count. It is generally sorted by frequency. Such a list is essential for many NLP tasks, including building language models, POS ta…
POSLinguistic Analysis of Sinhala YouTube Comments on Sinhala Music Videos: A Dataset Study
This research investigates the area of Music Information Retrieval (MIR) and Music Emotion Recognition (MER) in relation to Sinhala songs, an underexplored field in music studies. The purpose of this study is to analyze …
Emotion RecognitionInformation RetrievalMusic Emotion RecognitionMusic Information Retrieval+1