paper-with-me

Papers

MiLiC-Eval: Benchmarking Multilingual LLMs for China's Minority Languages

2025-03-03 · Chen Zhang, Mingxu Tao, Zhiyuan Liao, Yansong Feng

Large language models (LLMs) excel in high-resource languages but struggle with low-resource languages (LRLs), particularly those spoken by minority communities in China, such as Tibetan, Uyghur, Kazakh, and Mongolian. To systematically track the progress in these languages, we introduce MiLiC-Eval, a benchmark designed for minority languages in China, featuring 24K instances across 9 tasks. MiLiC-Eval focuses on underrepresented writing systems and provides a fine-grained assessment of linguistic and problem-solving skills. Our evaluation reveals that LLMs perform poorly on syntax-intensive tasks and multi-script languages. We further demonstrate how MiLiC-Eval can help advance LRL research in handling diverse writing systems and understanding the process of language adaptation.

📄 PDF Abstract BibTeX arXiv:2503.01150

Code (2)

luciusssss/milic-eval 공식 구현 pytorch
luciusssss/mlic-eval pytorch

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

How Chinese are Chinese Language Models? The Puzzling Lack of Language Policy in China's LLMs

2024-07-12 · Andrea W Wen-Yi, Unso Eun Seo Jo, Lu Jia Lin, David Mimno

Contemporary language models are increasingly multilingual, but Chinese LLM developers must navigate complex political and business considerations of language diversity. Language policy in China aims at influencing the p…

DiversityLanguage ModellingNavigate

Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs

2026-05-29 · Zhizhi Wang, Harini Suresh arxiv

LLMs increasingly answer questions about taxes, labor protections, healthcare, education, pensions, and administrative procedures, where usefulness often depends on the applicable jurisdiction. Multilingual users may wri…

TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking

2025-02-16 · Shahriar Kabir Nahin, Rabindra Nath Nandi, Sagor Sarker, Quazi Sarwar Muhtaseem 외

In this paper, we present TituLLMs, the first large pretrained Bangla LLMs, available in 1B and 3B parameter sizes. Due to computational constraints during both training and inference, we focused on smaller models. To tr…

Benchmarking

Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030

2025-05-12 · Mouxiao Bian, Rongzhao Zhang, Chao Ding, Xinwei Peng 외

Large Language Models (LLMs) are poised to transform healthcare under China's Healthy China 2030 initiative, yet they introduce new ethical and patient-safety challenges. We present a novel 12,000-item Q&A benchmark cove…

BenchmarkingEthics

MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Chatbots and Dialogue Evaluators

2025-05-28 · John Mendonça, Alon Lavie, Isabel Trancoso

As the capabilities of chatbots and their underlying LLMs continue to dramatically improve, evaluating their performance has increasingly become a major blocker to their further development. A major challenge is the avai…

BenchmarkingChatbotDialogue Evaluation