paper-with-me

Papers

Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking

2024-05-07 · Emre Can Acikgoz, Mete Erdogan, Deniz Yuret

Large Language Models (LLMs) are becoming crucial across various fields, emphasizing the urgency for high-quality models in underrepresented languages. This study explores the unique challenges faced by low-resource languages, such as data scarcity, model selection, evaluation, and computational limitations, with a special focus on Turkish. We conduct an in-depth analysis to evaluate the impact of training strategies, model choices, and data availability on the performance of LLMs designed for underrepresented languages. Our approach includes two methodologies: (i) adapting existing LLMs originally pretrained in English to understand Turkish, and (ii) developing a model from the ground up using Turkish pretraining data, both supplemented with supervised fine-tuning on a novel Turkish instruction-tuning dataset aimed at enhancing reasoning capabilities. The relative performance of these methods is evaluated through the creation of a new leaderboard for Turkish LLMs, featuring benchmarks that assess different reasoning and knowledge skills. Furthermore, we conducted experiments on data and model scaling, both during pretraining and fine-tuning, simultaneously emphasizing the capacity for knowledge transfer across languages and addressing the challenges of catastrophic forgetting encountered during fine-tuning on a different language. Our goal is to offer a detailed guide for advancing the LLM framework in low-resource linguistic contexts, thereby making natural language processing (NLP) benefits more globally accessible.

📄 PDF Abstract BibTeX arXiv:2405.04685

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingModel SelectionTransfer Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

BosphorusSign22k Sign Language Recognition Dataset

2020-04-02 · LREC 2020 5 · Oğulcan Özdemir, Ahmet Alp Kındıroğlu, Necati Cihan Camgöz, Lale Akarun

Sign Language Recognition is a challenging research domain. It has recently seen several advancements with the increased availability of data. In this paper, we introduce the BosphorusSign22k, a publicly available large …

Sign Language ProductionSign Language RecognitionVideo Recognition

BosphorusSign: A Turkish Sign Language Recognition Corpus in Health and Finance Domains

2016-05-01 · LREC 2016 5 · Necati Cihan Camg{\"o}z, Ahmet Alp K{\i}nd{\i}ro{\u{g}}lu, Serpil Karab{\"u}kl{\"u}, Meltem Kelepir 외

There are as many sign languages as there are deaf communities in the world. Linguists have been collecting corpora of different sign languages and annotating them extensively in order to study and understand their prope…

Sign Language Recognition

Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları

2025-08-18 · M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Banu Diri 외 arxiv

Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these models remains a challenge, particularly …

Emotion Recognition for Low-Resource Turkish: Fine-Tuning BERTurk on TREMO and Testing on Xenophobic Political Discourse

2025-05-17 · Darmawan Wicaksono, Hasri Akbar Awal Rozaq, Nevfel Boz

Social media platforms like X (formerly Twitter) play a crucial role in shaping public discourse and societal norms. This study examines the term Sessiz Istila (Silent Invasion) on Turkish social media, highlighting the …

Decision MakingEmotion RecognitionMarketingPublic Relations+1

Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish

2025-08-22 · Yakup Abrek Er, Ilker Kesen, Gözde Gül Şahin, Aykut Erdem arxiv

We introduce Cetvel, a comprehensive benchmark designed to evaluate large language models (LLMs) in Turkish. Existing Turkish benchmarks often lack either task diversity or culturally relevant content, or both. Cetvel ad…

Grammatical Error CorrectionMachine TranslationQuestion Answering