Stemming -- The Evolution and Current State with a Focus on Bangla
Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing step in language analysis, is essential for low-resource, highly-inflectional languages like Bangla, because it can reduce the complexity of algorithms and models by significantly reducing the number of words the algorithm needs to consider. This paper conducts a comprehensive survey of stemming approaches, emphasizing the importance of handling morphological variants effectively. While exploring the landscape of Bangla stemming, it becomes evident that there is a significant gap in the existing literature. The paper highlights the discontinuity from previous research and the scarcity of accessible implementations for replication. Furthermore, it critiques the evaluation methodologies, stressing the need for more relevant metrics. In the context of Bangla's rich morphology and diverse dialects, the paper acknowledges the challenges it poses. To address these challenges, the paper suggests directions for Bangla stemmer development. It concludes by advocating for robust Bangla stemmers and continued research in the field to enhance language analysis and processing.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
N-gram Statistical Stemmer for Bangla Corpus
Stemming is a process that can be utilized to trim inflected words to stem or root form. It is useful for enhancing the retrieval effectiveness, especially for text search in order to solve the mismatch problems. Previou…
ClusteringRetrievalBangla Parts-of-Speech Tagging using Bangla Stemmer and Rule based Analyzer
Parts-of-Speech (POS) tagging plays vital roles in the field of Natural Language Processing (NLP), such as - machine translation, spell checker, information retrieval, speech processing, emotion analysis and so on. Bangl…
Emotion RecognitionInformation RetrievalMachine TranslationPart-Of-Speech Tagging+4Advancing Bangla Machine Translation Through Informal Datasets
Bangla is the sixth most widely spoken language globally, with approximately 234 million native speakers. However, progress in open-source Bangla machine translation remains limited. Most online resources are in English …
Machine TranslationEvolution of exchange rate regime: Impact of macroeconomy of Bangladesh
Bangladesh has experienced two distinct exchange rate regimes: a fixed exchange rate system from January 1972 to May 2003 and a floating one since June 2003. After adopting the floating exchange rate regime, Bangladesh p…
A Study of fastText Word Embedding Effects in Document Classification in Bangla Language
Natural language processing is the current topic due to many important tasks like document classification, named entity recognition, opinion mining, sentiment analysis, textual entailment, etc. Such types of task in the …
ClassificationDocument ClassificationGeneral ClassificationLemmatization+6