paper-with-me

홈 › Papers

MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP

2025-06-04 · Kurt Micallef, Claudia Borg

Large Language Models (LLMs) have demonstrated remarkable performance across various Natural Language Processing (NLP) tasks, largely due to their generalisability and ability to perform tasks without additional training. However, their effectiveness for low-resource languages remains limited. In this study, we evaluate the performance of 55 publicly available LLMs on Maltese, a low-resource language, using a newly introduced benchmark covering 11 discriminative and generative tasks. Our experiments highlight that many models perform poorly, particularly on generative tasks, and that smaller fine-tuned models often perform better across all tasks. From our multidimensional analysis, we investigate various factors impacting performance. We conclude that prior exposure to Maltese during pre-training and instruction-tuning emerges as the most important factor. We also examine the trade-offs between fine-tuning and prompting, highlighting that while fine-tuning requires a higher initial cost, it yields better performance and lower inference costs. Through this work, we aim to highlight the need for more inclusive language technologies and recommend that researchers working with low-resource languages consider more "traditional" language modelling approaches.

📄 PDF Abstract BibTeX arXiv:2506.04385

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingLanguage Modelling

Similar Papers 제목 키워드 기반

Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT

2024-04-03 · Amirhossein Abaskohi, Sara Baruni, Mostafa Masoudi, Nesa Abbasi 외

This paper explores the efficacy of large language models (LLMs) for Persian. While ChatGPT and consequent LLMs have shown remarkable performance in English, their efficiency for more low-resource languages remains an op…

BenchmarkingGeneral KnowledgeMath

Benchmarking Floworks against OpenAI & Anthropic: A Novel Framework for Enhanced LLM Function Calling

2024-10-23 · Nirav Bhan, Shival Gupta, Sai Manaswini, Ritik Baba 외

Large Language Models (LLMs) have shown remarkable capabilities in various domains, yet their economic impact has been limited by challenges in tool use and function calling. This paper introduces ThorV2, a novel archite…

Benchmarking

Benchmarking LLAMA Model Security Against OWASP Top 10 For LLM Applications

2026-01-27 · Nourin Shahin, Izzat Alsmadi arxiv

As large language models (LLMs) move from research prototypes to enterprise systems, their security vulnerabilities pose serious risks to data privacy and system integrity. This study benchmarks various Llama model varia…

TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking

2025-02-16 · Shahriar Kabir Nahin, Rabindra Nath Nandi, Sagor Sarker, Quazi Sarwar Muhtaseem 외

In this paper, we present TituLLMs, the first large pretrained Bangla LLMs, available in 1B and 3B parameter sizes. Due to computational constraints during both training and inference, we focused on smaller models. To tr…

Benchmarking

STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator

2026-04-27 · Alessio Sordo, Lingxiao Du, Meeka-Hanna Lenisa, Evgeny Bogdanov 외 arxiv

The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific evaluation datasets; however, the collection of such datasets is challen…