Enhancing Assamese NLP Capabilities: Introducing a Centralized Dataset Repository
This paper introduces a centralized, open-source dataset repository designed to advance NLP and NMT for Assamese, a low-resource language. The repository, available at GitHub, supports various tasks like sentiment analysis, named entity recognition, and machine translation by providing both pre-training and fine-tuning corpora. We review existing datasets, highlighting the need for standardized resources in Assamese NLP, and discuss potential applications in AI-driven research, such as LLMs, OCR, and chatbots. While promising, challenges like data scarcity and linguistic diversity remain. The repository aims to foster collaboration and innovation, promoting Assamese language research in the digital age.
Code (1)
Tasks
DiversityMachine Translationnamed-entity-recognitionNamed Entity RecognitionNMTOptical Character Recognition (OCR)Sentiment AnalysisTranslationSimilar Papers 제목 키워드 기반
AsNER - Annotated Dataset and Baseline for Assamese Named Entity recognition
We present the AsNER, a named entity annotation dataset for low resource Assamese language with a baseline Assamese NER model. The dataset contains about 99k tokens comprised of text from the speech of the Prime Minister…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1AsNER -- Annotated Dataset and Baseline for Assamese Named Entity recognition
We present the AsNER, a named entity annotation dataset for low resource Assamese language with a baseline Assamese NER model. The dataset contains about 99k tokens comprised of text from the speech of the Prime Minister…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1Development of Assamese Rule based Stemmer using WordNet
Stemming is a technique that reduces any inflected word to its root form. Assamese is a morphologically rich, scheduled Indian language. There are various forms of suffixes applied to a word in various contexts. Such inf…
Image Caption Generation for Low-Resource Assamese Language
Image captioning is a prominent Artificial Intelligence (AI) research area that deals with visual recognition and a linguistic description of the image. It is an interdisciplinary field concerning how computers can see a…
Caption GenerationDecoderImage CaptioningMachine Translation+1Phonetic, Semantic, and Articulatory Features in Assamese-Bengali Cognate Detection
In this paper, we propose a method to detect if words in two similar languages, Assamese and Bengali, are cognates. We mix phonetic, semantic, and articulatory features and use the cognate detection task to analyze the r…