Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
Nepali, a low-resource language, faces significant challenges in building an effective information retrieval system due to the unavailability of annotated data and computational linguistic resources. In this study, we attempt to address this gap by preparing a pair-structured Nepali Question-Answer dataset. We focus on Frequently Asked Questions (FAQs) for passport-related services, building a data set for training and evaluation of IR models. In our study, we have fine-tuned transformer-based embedding models for semantic similarity in question-answer retrieval. The fine-tuned models were compared with the baseline BM25. In addition, we implement a hybrid retrieval approach, integrating fine-tuned models with BM25, and evaluate the performance of the hybrid retrieval. Our results show that the fine-tuned SBERT-based models outperform BM25, whereas multilingual E5 embedding-based models achieve the highest retrieval performance among all evaluated models.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalSemantic SimilarityQuestion AnsweringSimilar Papers 제목 키워드 기반
Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering
Legal domains in high-resource languages like English have widely adopted artificial intelligence for legal question answering. However, data scarcity in low resource languages such as Nepali has limited the training of …
Question AnsweringAnswer GenerationNepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
Modern Translation Systems heavily rely on high-quality, large parallel datasets for state-of-the-art performance. However, such resources are largely unavailable for most of the South Asian languages. Among them, Nepali…
Machine TranslationAbstractive Summarization of Low resourced Nepali language using Multilingual Transformers
Automatic text summarization in Nepali language is an unexplored area in natural language processing (NLP). Although considerable research has been dedicated to extractive summarization, the area of abstractive summariza…
Abstractive Text SummarizationArticlesExtractive SummarizationHeadline Generation+2Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
Nepali, a low-resource language spoken by over 32 million people, continues to face challenges in natural language processing (NLP) due to its complex grammar, agglutinative morphology, and limited availability of high-q…
Text GenerationGenerative AI for Named Entity Recognition in Low-Resource Language Nepali
Generative Artificial Intelligence (GenAI), particularly Large Language Models (LLMs), has significantly advanced Natural Language Processing (NLP) tasks, such as Named Entity Recognition (NER), which involves identifyin…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER