paper-with-me

Papers

NusaBERT: Teaching IndoBERT to be Multilingual and Multicultural

2024-03-04 · Wilson Wongso, David Samuel Setiawan, Steven Limcorn, Ananto Joyoadikusumo

Indonesia's linguistic landscape is remarkably diverse, encompassing over 700 languages and dialects, making it one of the world's most linguistically rich nations. This diversity, coupled with the widespread practice of code-switching and the presence of low-resource regional languages, presents unique challenges for modern pre-trained language models. In response to these challenges, we developed NusaBERT, building upon IndoBERT by incorporating vocabulary expansion and leveraging a diverse multilingual corpus that includes regional languages and dialects. Through rigorous evaluation across a range of benchmarks, NusaBERT demonstrates state-of-the-art performance in tasks involving multiple languages of Indonesia, paving the way for future natural language understanding research for under-represented languages.

📄 PDF Abstract BibTeX arXiv:2403.01817

Code (1)

LazarusNLP/NusaBERT 공식 구현 pytorch

Tasks

DiversityNatural Language Understanding

Similar Papers 제목 키워드 기반

Enhancing Multilingual Information Retrieval in Mixed Human Resources Environments: A RAG Model Implementation for Multicultural Enterprise

2024-01-03 · Syed Rameel Ahmad

The advent of Large Language Models has revolutionized information retrieval, ushering in a new era of expansive knowledge accessibility. While these models excel in providing open-world knowledge, effectively extracting…

Information RetrievalRAGRetrievalRetrieval-augmented Generation+1

YEZE at SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization via Heterogeneous Ensembling

2026-05-07 · Fengze Guo, Yue Chang arxiv

This paper presents our system for SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization, which identifies polarized social media content in 22 languages through three subtasks: bi…

Multi-Task LearningData Augmentation

CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks

2026-04-21 · Peiqin Lin, Chenyang Lyu, Wenjiang Luo, Haotian Ye 외 arxiv

Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities. However, these benchmarks prioritize generic language understanding or…

Leveraging LaBSE with Progressive Curriculum Learning for Multicultural Polarization

2026-06-19 · Sachin Sundar, Sandeep Kumar, Mothish M arxiv

Detecting online polarization remains a critical challenge, particularly in multilingual and multicultural contexts where intergroup hostility is prevalent. The problem is particularly challenging due to the data scarcit…

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

2024-10-16 · Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha 외

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we …

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)