Bangla Parts-of-Speech Tagging using Bangla Stemmer and Rule based Analyzer
Parts-of-Speech (POS) tagging plays vital roles in the field of Natural Language Processing (NLP), such as - machine translation, spell checker, information retrieval, speech processing, emotion analysis and so on. Bangla is a very inflectional language that induces many variants from a single word. Although there is a few POS Tagger in Bangla language,very small of them address the essence of suffices to identify tag of the words. In this regard, we propose an automated POS Tagging system for Bangla language based on word-suffixes. In our system, we use our own stemming technique to retrieve a possible minimum root words and apply rules according to different forms of suffixes. Moreover, we incorporate a Bangla vocabulary that contains more than 45,000 words with their default tag and a patterned based verb-data-set. These facilitate to improve tagging efficiency of Bangla POS Tagger. We experiment our proposed system on a Bangla text corpus. The result shows that our proposed Bangla POS Tagger has outperformed the known related tagging systems.
Code (1)
Tasks
Emotion RecognitionInformation RetrievalMachine TranslationPart-Of-Speech TaggingPOSPOS TaggingRetrievalTAGSimilar Papers 제목 키워드 기반
Stemming -- The Evolution and Current State with a Focus on Bangla
Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing s…
Bangla Word Clustering Based on Tri-gram, 4-gram and 5-gram Language Model
In this paper, we describe a research method that generates Bangla word clusters on the basis of relating to meaning in language and contextual similarity. The importance of word clustering is in parts of speech (POS) ta…
ClusteringLanguage ModelingLanguage ModellingPOS+5Bangla Natural Language Processing: A Comprehensive Analysis of Classical, Machine Learning, and Deep Learning Based Methods
The Bangla language is the seventh most spoken language, with 265 million native and non-native speakers worldwide. However, English is the predominant language for online resources and technical knowledge, journals, and…
ArticlesBIG-bench Machine LearningMachine Translationnamed-entity-recognition+10BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla
Medical question answering (QA) systems have become crucial tools for providing reliable health information. But they remain very unexplored for low-resource languages like Bangla due to limited datasets and systems tail…
Part-Of-Speech TaggingQuestion AnsweringByakto Speech: Real-time long speech synthesis with convolutional neural network: Transfer learning from English to Bangla
Speech synthesis is one of the challenging tasks to automate by deep learning, also being a low-resource language there are very few attempts at Bangla speech synthesis. Most of the existing works can't work with anythin…
Deep Learningspeech-recognitionSpeech RecognitionSpeech Synthesis+1