paper-with-me

Papers

Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation

2024-12-31 · M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Banu Diri, Savaş Yıldırım, Öner Aytaş

Language models have made remarkable advancements in understanding and generating human language, achieving notable success across a wide array of applications. However, evaluating these models remains a significant challenge, particularly for resource-limited languages such as Turkish. To address this gap, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is constructed from a carefully curated dataset comprising 6200 multiple-choice questions across 62 sections, selected from a pool of 280000 questions spanning 67 disciplines and over 800 topics within the Turkish education system. This benchmark provides a transparent, reproducible, and culturally relevant tool for evaluating model performance. It serves as a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text and fostering the development of more robust and accurate language models. In this study, we evaluate state-of-the-art LLMs on TR-MMLU, providing insights into their strengths and limitations for Turkish-specific tasks. Our findings reveal critical challenges, such as the impact of tokenization and fine-tuning strategies, and highlight areas for improvement in model design. By setting a new standard for evaluating Turkish language models, TR-MMLU aims to inspire future innovations and support the advancement of Turkish NLP research.

📄 PDF Abstract BibTeX arXiv:2501.00593

Code (0)

등록된 구현이 없습니다.

Tasks

Language Model EvaluationLanguage ModelingLanguage ModellingLarge Language ModelMMLUMultiple-choice

Similar Papers 제목 키워드 기반

Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi

2025-08-18 · M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Sercan Karakaş 외 arxiv

Tokenization is a fundamental preprocessing step in Natural Language Processing (NLP), significantly impacting the capability of large language models (LLMs) to capture linguistic and semantic nuances. This study introdu…

Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları

2025-08-18 · M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Banu Diri 외 arxiv

Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these models remains a challenge, particularly …

TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish

2024-07-17 · Arda Yüksel, Abdullatif Köksal, Lütfi Kerem Şenel, Anna Korhonen 외

Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of Large Language Models (LLMs). While existing benchmarks employ automatic translation for multilingual evaluati…

MathMultiple-choiceQuestion Answering

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

2025-02-10 · M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Sercan Karakaş 외

Tokenization is a fundamental preprocessing step in NLP, directly impacting large language models' (LLMs) ability to capture syntactic, morphosyntactic, and semantic structures. This paper introduces a novel framework fo…

MMLUMorphological AnalysisMultiple-choicevalid

A Turkish Educational Crossword Puzzle Generator

2024-05-11 · Kamyar Zeinalipour, Yusuf Gökberk Keptiğ, Marco Maggini, Leonardo Rigutini 외

This paper introduces the first Turkish crossword puzzle generator designed to leverage the capabilities of large language models (LLMs) for educational purposes. In this work, we introduced two specially created dataset…