paper-with-me

Papers

N-gram Statistical Stemmer for Bangla Corpus

2019-12-25 · Rabeya Sadia, Md. Ataur Rahman, Md. Hanif Seddiqui

Stemming is a process that can be utilized to trim inflected words to stem or root form. It is useful for enhancing the retrieval effectiveness, especially for text search in order to solve the mismatch problems. Previous research on Bangla stemming mostly relied on eliminating multiple suffixes from a solitary word through a recursive rule based procedure to recover progressively applicable relative root. Our proposed system has enhanced the aforementioned exploration by actualizing one of the stemming algorithms called N-gram stemming. By utilizing an affiliation measure called dice coefficient, related sets of words are clustered depending on their character structure. The smallest word in one cluster may be considered as the stem. We additionally analyzed Affinity Propagation clustering algorithms with coefficient similarity as well as with median similarity. Our result indicates N-gram stemming techniques to be effective in general which gave us around 87% accurate clusters.

📄 PDF Abstract BibTeX arXiv:1912.11612

Code (1)

shaoncsecu/Bangla_n-gram_Stemmer 공식 구현

Tasks

ClusteringRetrieval

Similar Papers 제목 키워드 기반

Bangla Parts-of-Speech Tagging using Bangla Stemmer and Rule based Analyzer

2016-06-09 · 18th International Conference on Computer and Information Technology (ICCIT) 2016 6 · Md. Nesarul Hoque, Md. Hanif Seddiqui

Parts-of-Speech (POS) tagging plays vital roles in the field of Natural Language Processing (NLP), such as - machine translation, spell checker, information retrieval, speech processing, emotion analysis and so on. Bangl…

Emotion RecognitionInformation RetrievalMachine TranslationPart-Of-Speech Tagging+4

Stemming -- The Evolution and Current State with a Focus on Bangla

2025-08-21 · Abhijit Paul, Mashiat Amin Farin, Sharif Md. Abdullah, Ahmedul Kabir 외 arxiv

Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing s…

An Annotated Dataset and Automatic Approaches for Discourse Mode Identification in Low-resource Bengali Language

2022-07-01 · NAACL (MIA) 2022 7 · Salim Sazzed

The modes of discourse aid in comprehending the convention and purpose of various forms of languages used during communication. In this study, we introduce a discourse mode annotated corpus for the low-resource Bangla (a…

DescriptiveLanguage ModelingLanguage ModellingSentence

VAIYAKARANA : A Benchmark for Automatic Grammar Correction in Bangla

2024-06-20 · Pramit Bhattacharyya, Arnab Bhattacharya

Bangla (Bengali) is the fifth most spoken language globally and, yet, the problem of automatic grammar correction in Bangla is still in its nascent stage. This is mostly due to the need for a large corpus of grammaticall…

Sentence

BiST: A Gold Standard Bangla-English Bilingual Corpus for Sentence Structure and Tense Classification with Inter-Annotator Agreement

2026-04-06 · Abdullah Al Shafi, Swapnil Kundu Argha, M. A. Moyeen, Abdul Muntakim 외 arxiv

High-quality bilingual resources remain a critical bottleneck for advancing multilingual NLP in low-resource settings, particularly for Bangla. To mitigate this gap, we introduce BiST, a rigorously curated Bangla-English…

Representation LearningLanguage IdentificationText Generation