paper-with-me

Papers

NativQA: Multilingual Culturally-Aligned Natural Query for LLMs

2024-07-13 · Md. Arid Hasan, Maram Hasanain, Fatema Ahmad, Sahinur Rahman Laskar, Sunaya Upadhyay, Vrunda N Sukhadia, Mucahid Kutlu, Shammur Absar Chowdhury, Firoj Alam

Natural Question Answering (QA) datasets play a crucial role in evaluating the capabilities of large language models (LLMs), ensuring their effectiveness in real-world applications. Despite the numerous QA datasets that have been developed, there is a notable lack of region-specific datasets generated by native users in their own languages. This gap hinders the effective benchmarking of LLMs for regional and cultural specificities. Furthermore, it also limits the development of fine-tuned models. In this study, we propose a scalable, language-independent framework, NativQA, to seamlessly construct culturally and regionally aligned QA datasets in native languages, for LLM evaluation and tuning. We demonstrate the efficacy of the proposed framework by designing a multilingual natural QA dataset, \mnqa, consisting of ~64k manually annotated QA pairs in seven languages, ranging from high to extremely low resource, based on queries from native speakers from 9 regions covering 18 topics. We benchmark open- and closed-source LLMs with the MultiNativQA dataset. We also showcase the framework efficacy in constructing fine-tuning data especially for low-resource and dialectally-rich languages. We made both the framework NativQA and MultiNativQA dataset publicly available for the community (https://nativqa.gitlab.io).

📄 PDF Abstract BibTeX arXiv:2407.09823

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingQuestion Answering

Similar Papers 제목 키워드 기반

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs

2025-05-25 · Firoj Alam, Md Arid Hasan, Shammur Absar Chowdhury

Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and tasks. However, benchmarking their capabilities with multilingual spoken queries remains largely unexplored. In this st…

BenchmarkingDiversityQuestion Answering

CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAG

2026-04-28 · Nayeon Lee, Jiwoo Song, Byeongcheol Kang arxiv

Multilingual retrieval-augmented generation (mRAG) is often implemented within a fixed retrieval space, typically via query or document translation or multilingual embedding vector representations. However, this approach…

CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications

2025-08-03 · Raviraj Joshi, Rakesh Paul, Kanishk Singla, Anusha Kamath 외 arxiv

The increasing use of Large Language Models (LLMs) in agentic applications highlights the need for robust safety guard models. While content safety in English is well-studied, non-English languages lack similar advanceme…

Synthetic Data GenerationZero-shot GeneralizationCross-Lingual TransferMachine Translation

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

2026-05-01 · Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Yu-Gang Jiang 외 arxiv

As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety in diverse regulatory and cultural environments has become a critical challenge. However, existing multilingual benc…

Machine Translation

Disentangling Language and Culture for Evaluating Multilingual Large Language Models

2025-05-30 · Jiahao Ying, Wei Tang, Yiran Zhao, Yixin Cao 외

This paper introduces a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of LLMs. By decomposing the evaluation along the dimensions of linguistic medium and cultural context, this framew…