paper-with-me

Papers

The Impact of Model Scaling on Seen and Unseen Language Performance

2025-01-10 · Rhitabrat Pokharel, Sina Bagheri Nezhad, Ameeta Agrawal, Suresh Singh

The rapid advancement of Large Language Models (LLMs), particularly those trained on multilingual corpora, has intensified the need for a deeper understanding of their performance across a diverse range of languages and model sizes. Our research addresses this critical need by studying the performance and scaling behavior of multilingual LLMs in text classification and machine translation tasks across 204 languages. We systematically examine both seen and unseen languages across three model families of varying sizes in zero-shot and few-shot settings. Our findings show significant differences in scaling behavior between zero-shot and two-shot scenarios, with striking disparities in performance between seen and unseen languages. Model scale has little effect on zero-shot performance, which remains mostly flat. However, in two-shot settings, larger models show clear linear improvements in multilingual text classification. For translation tasks, however, only the instruction-tuned model showed clear benefits from scaling. Our analysis also suggests that overall resource levels, not just the proportions of pretraining languages, are better predictors of model performance, shedding light on what drives multilingual LLM effectiveness.

📄 PDF Abstract BibTeX arXiv:2501.05629

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationMultilingual text classificationtext-classificationText ClassificationTranslation

Similar Papers 제목 키워드 기반

Florenz: Scaling Laws for Systematic Generalization in Vision-Language Models

2025-03-12 · Julian Spravil, Sebastian Houben, Sven Behnke

Cross-lingual transfer enables vision-language models (VLMs) to perform vision tasks in various languages with training data only in one language. Current approaches rely on large pre-trained multilingual language models…

Cross-Lingual TransferImage CaptioningLarge Language ModelMachine Translation+3

Re-Evaluating the Impact of Unseen-Class Unlabeled Data on Semi-Supervised Learning Model

2025-03-02 · Rundong He, Yicong Dong, LanZhe Guo, Yilong Yin 외

Semi-supervised learning (SSL) effectively leverages unlabeled data and has been proven successful across various fields. Current safe SSL methods believe that unseen classes in unlabeled data harm the performance of SSL…

Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1

2025-08-13 · Petr Spelda, Vit Stritecky arxiv

Evaluation of reasoning language models gained importance after it was observed that they can combine their existing capabilities into novel traces of intermediate steps before task completion and that the traces can som…

Beyond Text Compression: Evaluating Tokenizers Across Scales

2025-06-03 · Jonas F. Lotz, António V. Lopes, Stephan Peitz, Hendra Setiawan 외

The choice of tokenizer can profoundly impact language model performance, yet accessible and reliable evaluations of tokenizer quality remain an open challenge. Inspired by scaling consistency, we show that smaller model…

Language ModelingLanguage ModellingText Compression

ZeroPrompt: Scaling Prompt-Based Pretraining to 1,000 Tasks Improves Zero-Shot Generalization

2022-01-18 · Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao 외

We propose a multitask pretraining approach ZeroPrompt for zero-shot generalization, focusing on task scaling and zero-shot prompting. While previous models are trained on only a few dozen tasks, we scale to 1,000 tasks …

Zero-shot GeneralizationZero-Shot Learning