paper-with-me

홈 › Papers

BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In this paper, we introduce 'BanglaBERT', a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect 27.5 GB of Bangla pretraining data (dubbed 'Bangla2B+') by crawling 110 popular Bangla sites. We introduce a new downstream task dataset on Natural Language Inference (NLI) and benchmark on four diverse NLU tasks covering text classification, sequence labeling, and span prediction. In the process, we bring them under the first-ever Bangla Language Understanding Evaluation (BLUE) benchmark. BanglaBERT achieves state-of-the-art results outperforming multilingual and monolingual models. We will make the BanglaBERT model, the new datasets, and a leaderboard publicly available to advance Bangla NLP.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingNatural Language InferenceNatural Language Understandingtext-classificationText Classification

Similar Papers 제목 키워드 기반

BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla

2021-01-01 · Findings (NAACL) 2022 7 · Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Kazi Samin 외

In this work, we introduce BanglaBERT, a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect 27.…

Document ClassificationLanguage ModelingLanguage ModellingNatural Language Inference+5

Bangla Hate Speech Classification with Fine-tuned Transformer Models

2025-12-02 · Yalda Keivan Jafari, Krishno Dey arxiv

Hate speech recognition in low-resource languages remains a difficult problem due to insufficient datasets, orthographic heterogeneity, and linguistic variety. Bangla is spoken by more than 230 million people of Banglade…

Hate Speech DetectionSpeech Recognition

LowResource at BLP-2023 Task 2: Leveraging BanglaBert for Low Resource Sentiment Analysis of Bangla Language

2023-11-21 · Aunabil Chakma, Masum Hasan

This paper describes the system of the LowResource Team for Task 2 of BLP-2023, which involves conducting sentiment analysis on a dataset composed of public posts and comments from diverse social media platforms. Our pri…

Sentiment AnalysisTask 2

LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target

2025-10-02 · Md Arid Hasan, Firoj Alam, Md Fahad Hossain, Usman Naseem 외 arxiv

Online social media platforms are central to everyday communication and information seeking. While these platforms serve positive purposes, they also provide fertile ground for the spread of hate speech, offensive langua…

Hate Speech Detection

AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African Languages

2021-09-10 · EMNLP 2021 11 · Machel Reid, Junjie Hu, Graham Neubig, Yutaka Matsuo

Reproducible benchmarks are crucial in driving progress of machine translation research. However, existing machine translation benchmarks have been mostly limited to high-resource or well-represented languages. Despite a…

Cross-Lingual TransferData AugmentationMachine TranslationTranslation