paper-with-me

홈 › Papers

Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO

2024-12-17 · Umer Butt, Stalin Varanasi, Günter Neumann

As the Information Retrieval (IR) field increasingly recognizes the importance of inclusivity, addressing the needs of low-resource languages remains a significant challenge. This paper introduces the first large-scale Urdu IR dataset, created by translating the MS MARCO dataset through machine translation. We establish baseline results through zero-shot learning for IR in Urdu and subsequently apply the mMARCO multilingual IR methodology to this newly translated dataset. Our findings demonstrate that the fine-tuned model (Urdu-mT5-mMARCO) achieves a Mean Reciprocal Rank (MRR@10) of 0.247 and a Recall@10 of 0.439, representing significant improvements over zero-shot results and showing the potential for expanding IR access for Urdu speakers. By bridging access gaps for speakers of low-resource languages, this work not only advances multilingual IR research but also emphasizes the ethical and societal importance of inclusive IR technologies. This work provides valuable insights into the challenges and solutions for improving language representation and lays the groundwork for future research, especially in South Asian languages, which can benefit from the adaptable methods used in this study.

📄 PDF Abstract BibTeX arXiv:2412.12997

Code (1)

UmerTariq1/Urdu_MsMarco_Translation_Retrieval 공식 구현 pytorch

Tasks

Information RetrievalMachine TranslationRetrievalZero-Shot Learning

Similar Papers 제목 키워드 기반

Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language

2025-08-03 · Jaskaranjeet Singh, Rakesh Thakur arxiv

Despite rapid advances in large language models (LLMs), low-resource languages remain excluded from NLP, limiting digital access for millions. We present PunGPT2, the first fully open-source Punjabi generative model suit…

Question Answering

KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation

2025-07-04 · Antoine Nzeyimana, Andre Niyongabo Rubungo arxiv

The recent mainstream adoption of large language model (LLM) technology is enabling novel applications in the form of chatbots and virtual assistants across many domains. With the aim of grounding LLMs in trusted domains…

Khana: A Comprehensive Indian Cuisine Dataset

2025-09-07 · Omkar Prabhu arxiv

As global interest in diverse culinary experiences grows, food image models are essential for improving food-related applications by enabling accurate food recognition, recipe suggestions, dietary tracking, and automated…

Image Classification

MME-RAG: Multi-Manager-Expert Retrieval-Augmented Generation for Fine-Grained Entity Recognition in Task-Oriented Dialogues

2025-11-15 · Liang Xue, Haoyu Liu, Yajun Tian, Xinyu Zhong 외 arxiv

Fine-grained entity recognition is crucial for reasoning and decision-making in task-oriented dialogues, yet current large language models (LLMs) continue to face challenges in domain adaptation and retrieval controllabi…

Domain GeneralizationDomain Adaptation

Bridging Dual Knowledge Graphs for Multi-Hop Question Answering in Construction Safety

2025-07-18 · Yuxin Zhang, Xi Wang, Mo Hu, Zhenyu Zhang arxiv

Information retrieval and question answering from safety regulations are essential for automated construction compliance checking but are hindered by the linguistic and structural complexity of regulatory text. Many quer…

Multi-hop Question AnsweringInformation RetrievalKnowledge Graphs