paper-with-me

Papers

LLM Probe: Evaluating LLMs for Low-Resource Languages

2026-03-31 · Hailay Kidu Teklehaymanot, Gebrearegawi Gebremariam, Wolfgang Nejdl arxiv

Despite rapid advances in large language models (LLMs), their linguistic abilities in low-resource and morphologically rich languages are still not well understood due to limited annotated resources and the absence of standardized evaluation frameworks. This paper presents LLM Probe, a lexicon-based assessment framework designed to systematically evaluate the linguistic skills of LLMs in low-resource language environments. The framework analyzes models across four areas of language understanding: lexical alignment, part-of-speech recognition, morphosyntactic probing, and translation accuracy. To illustrate the framework, we create a manually annotated benchmark dataset using a low-resource Semitic language as a case study. The dataset comprises bilingual lexicons with linguistic annotations, including part-of-speech tags, grammatical gender, and morphosyntactic features, which demonstrate high inter-annotator agreement to ensure reliable annotations. We test a variety of models, including causal language models and sequence-to-sequence architectures. The results reveal notable differences in performance across various linguistic tasks: sequence-to-sequence models generally excel in morphosyntactic analysis and translation quality, whereas causal models demonstrate strong performance in lexical alignment but exhibit weaker translation accuracy. Our results emphasize the need for linguistically grounded evaluation to better understand LLM limitations in low-resource settings. We release LLM Probe and the accompanying benchmark dataset as open-source tools to promote reproducible benchmarking and to support the development of more inclusive multilingual language technologies.

📄 PDF Abstract BibTeX arXiv:2603.29517

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

SemBench: A Universal Semantic Framework for LLM Evaluation

2026-03-12 · Mikel Zubillaga, Naiara Perez, Oscar Sainz, German Rigau arxiv

Recent progress in Natural Language Processing (NLP) has been driven by the emergence of Large Language Models (LLMs), which exhibit remarkable generative and reasoning capabilities. However, despite their success, evalu…

MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language

2025-05-20 · Seyoung Song, Seogyeong Jeong, Eunsu Kim, Jiho Jin 외

Evaluating text generation capabilities of large language models (LLMs) is challenging, particularly for low-resource languages where methods for direct assessment are scarce. We propose MUG-Eval, a novel framework that …

Text Generation

High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models

2024-02-19 · Michela Lorandi, Anya Belz

The performance of NLP methods for severely under-resourced languages cannot currently hope to match the state of the art in NLP methods for well resourced languages. We explore the extent to which pretrained large langu…

Data-to-Text GenerationText Generation

Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages

2024-04-17 · Zihao Li, Yucheng Shi, Zirui Liu, Fan Yang 외

The development of Large Language Models (LLMs) relies on extensive text corpora, which are often unevenly distributed across languages. This imbalance results in LLMs performing significantly better on high-resource lan…

SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?

2025-06-05 · Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry 외

Evaluating machine translation (MT) quality for under-resourced African languages remains a significant challenge, as existing metrics often suffer from limited language coverage and poor performance in low-resource sett…

Machine TranslationSentence