paper-with-me

홈 › Papers

Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications

2026-03-04 · Funghang Limbu Begha, Praveen Acharya, Bal Krishna Bal arxiv

Nepali, a low-resource language, faces significant challenges in building an effective information retrieval system due to the unavailability of annotated data and computational linguistic resources. In this study, we attempt to address this gap by preparing a pair-structured Nepali Question-Answer dataset. We focus on Frequently Asked Questions (FAQs) for passport-related services, building a data set for training and evaluation of IR models. In our study, we have fine-tuned transformer-based embedding models for semantic similarity in question-answer retrieval. The fine-tuned models were compared with the baseline BM25. In addition, we implement a hybrid retrieval approach, integrating fine-tuned models with BM25, and evaluate the performance of the hybrid retrieval. Our results show that the fine-tuned SBERT-based models outperform BM25, whereas multilingual E5 embedding-based models achieve the highest retrieval performance among all evaluated models.

📄 PDF Abstract BibTeX arXiv:2603.13320

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalSemantic SimilarityQuestion Answering

Similar Papers 제목 키워드 기반

Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering

2026-04-20 · Samir Wagle, Abiral Adhikari, Reewaj Khanal, Batsal Bhandari 외 arxiv

Legal domains in high-resource languages like English have widely adopted artificial intelligence for legal question answering. However, data scarcity in low resource languages such as Nepali has limited the training of …

Question AnsweringAnswer Generation

NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments

2026-03-14 · Rupak Raj Ghimire, Bipesh Subedi, Balaram Prasain, Prakash Poudyal 외 arxiv

Modern Translation Systems heavily rely on high-quality, large parallel datasets for state-of-the-art performance. However, such resources are largely unavailable for most of the South Asian languages. Among them, Nepali…

Machine Translation

Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers

2024-09-29 · Prakash Dhakal, Daya Sagar Baral

Automatic text summarization in Nepali language is an unexplored area in natural language processing (NLP). Although considerable research has been dedicated to extractive summarization, the area of abstractive summariza…

Abstractive Text SummarizationArticlesExtractive SummarizationHeadline Generation+2

Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer

2025-12-16 · Adarsha Shrestha, Basanta Pokharel, Binit Shrestha, Smriti Adhikari 외 arxiv

Nepali, a low-resource language spoken by over 32 million people, continues to face challenges in natural language processing (NLP) due to its complex grammar, agglutinative morphology, and limited availability of high-q…

Text Generation

Generative AI for Named Entity Recognition in Low-Resource Language Nepali

2025-03-12 · Sameer Neupane, Jeevan Chapagain, Nobal B. Niraula, Diwa Koirala

Generative Artificial Intelligence (GenAI), particularly Large Language Models (LLMs), has significantly advanced Natural Language Processing (NLP) tasks, such as Named Entity Recognition (NER), which involves identifyin…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER