paper-with-me

Papers

From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge

2025-10-22 · Nafis Chowdhury, Moinul Haque, Anika Ahmed, Nazia Tasnim, Md. Istiak Hossain Shihab, Sajjadur Rahman, Farig Sadeque arxiv

Recent progress in NLP research has demonstrated remarkable capabilities of large language models (LLMs) across a wide range of tasks. While recent multilingual benchmarks have advanced cultural evaluation for LLMs, critical gaps remain in capturing the nuances of low-resource cultures. Our work addresses these limitations through a Bengali Language Cultural Knowledge (BLanCK) dataset including folk traditions, culinary arts, and regional dialects. Our investigation of several multilingual language models shows that while these models perform well in non-cultural categories, they struggle significantly with cultural knowledge and performance improves substantially across all models when context is provided, emphasizing context-aware architectures and culturally curated training data.

📄 PDF Abstract BibTeX arXiv:2510.20043

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Systematic Study and Analysis of Bengali Folklore with Natural Language Processing Systems

2022-03-13 · Mustain Billah, Md. Mynoddin, Mostafijur Rahman Akhond, Md. Nasim Adnan 외

Folklore, a solid branch of folk literature, is the hallmark of any nation or any society. Such as oral tradition; as proverbs or jokes, it also includes material culture as well as traditional folk beliefs, and various …

Cultural Vocal Bursts Intensity Prediction

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

2026-03-22 · Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta, Rubaya Tabassum 외 arxiv

Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented in multimodal evaluation. To address this gap, we introduce BanglaVers…

Visual Question AnsweringVisual Grounding

Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE

2025-11-07 · Firoj Ahmmed Patwary, Abdullah Al Noman arxiv

Tokenization is an important first step in Natural Language Processing (NLP) pipelines because it decides how models learn and represent linguistic information. However, current subword tokenizers like SentencePiece or H…

News Classification

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

2026-01-30 · Xinyu Liu, Zixuan Xie, Amir Moeini, Claire Chen 외 arxiv

While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models (LLMs), the absence of many folklore lemmas in Mathlib remains a persistent b…

Mathematical Reasoning

BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali

2025-11-25 · Abdullah Al Sefat arxiv

Large language models excel on broad multilingual benchmarks but remain to be evaluated extensively in figurative and culturally grounded reasoning, especially in low-resource contexts. We present BengaliFig, a compact y…