paper-with-me

홈 › Papers

Reveal-Bangla: A Dataset for Cross-Lingual Multi-Step Reasoning Evaluation

2025-08-12 · Khondoker Ittehadul Islam, Gabriele Sarti arxiv

Language models have demonstrated remarkable performance on complex multi-step reasoning tasks. However, their evaluation has been predominantly confined to high-resource languages such as English. In this paper, we introduce a manually translated Bangla multi-step reasoning dataset derived from the English Reveal dataset, featuring both binary and non-binary question types. We conduct a controlled evaluation of English-centric and Bangla-centric multilingual small language models on the original dataset and our translated version to compare their ability to exploit relevant reasoning steps to produce correct answers. Our results show that, in comparable settings, reasoning context is beneficial for more challenging non-binary questions, but models struggle to employ relevant Bangla reasoning steps effectively. We conclude by exploring how reasoning steps contribute to models' predictions, highlighting different trends across models and languages.

📄 PDF Abstract BibTeX arXiv:2508.08933

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Evaluation of Bangla Word Analogies

2023-04-10 · Mousumi Akter, Souvika Sarkar, Shubhra Kanti Karmaker

This paper presents a high-quality dataset for evaluating the quality of Bangla word embeddings, which is a fundamental task in the field of Natural Language Processing (NLP). Despite being the 7th most-spoken language i…

BenchmarkingWord Embeddings

Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language

2025-10-10 · Adity Khisa, Nusrat Jahan Lia, Tasnim Mahfuz Nafis, Zarif Masud 외 arxiv

As an Indo-Aryan language with limited available data, Chakma remains largely underrepresented in language models. In this work, we introduce a novel corpus of contextually coherent Bangla-transliterated Chakma, curated …

Transfer Learning

Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement

2025-08-27 · Mohammed Rakibul Hasan, Rafi Majid, Ahanaf Tahmid arxiv

In this paper, we introduce Bangla-Bayanno, an open-ended Visual Question Answering (VQA) Dataset in Bangla, a widely used, low-resource language in multimodal AI research. The majority of existing datasets are either ma…

Visual Question Answering

Crosslingual Retrieval Augmented In-context Learning for Bangla

2023-11-01 · Xiaoqian Li, Ercong Nie, Sheng Liang

The promise of Large Language Models (LLMs) in Natural Language Processing has often been overshadowed by their limited performance in low-resource languages such as Bangla. To address this, our paper presents a pioneeri…

In-Context LearningRetrieval

BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In this paper, we introduce 'BanglaBERT', a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect …

Language ModelingLanguage ModellingNatural Language InferenceNatural Language Understanding+2