paper-with-me

홈 › Papers

Why We Build Local Large Language Models: An Observational Analysis from 35 Japanese and Multilingual LLMs

2024-12-19 · Koshiro Saito, Sakae Mizuki, Masanari Ohi, Taishi Nakamura, Taihei Shiotani, Koki Maeda, Youmi Ma, Kakeru Hattori, Kazuki Fujii, Takumi Okamoto, Shigeki Ishida, Hiroya Takamura, Rio Yokota, Naoaki Okazaki

Why do we build local large language models (LLMs)? What should a local LLM learn from the target language? Which abilities can be transferred from other languages? Do language-specific scaling laws exist? To explore these research questions, we evaluated 35 Japanese, English, and multilingual LLMs on 19 evaluation benchmarks for Japanese and English, taking Japanese as a local language. Adopting an observational approach, we analyzed correlations of benchmark scores, and conducted principal component analysis (PCA) on the scores to derive \textit{ability factors} of local LLMs. We found that training on English text can improve the scores of academic subjects in Japanese (JMMLU). In addition, it is unnecessary to specifically train on Japanese text to enhance abilities for solving Japanese code generation, arithmetic reasoning, commonsense, and reading comprehension tasks. In contrast, training on Japanese text could improve question-answering tasks about Japanese knowledge and English-Japanese translation, which indicates that abilities for solving these two tasks can be regarded as \textit{Japanese abilities} for LLMs. Furthermore, we confirmed that the Japanese abilities scale with the computational budget for Japanese text.

📄 PDF Abstract BibTeX arXiv:2412.14471

Code (0)

등록된 구현이 없습니다.

Tasks

Arithmetic ReasoningCode GenerationQuestion AnsweringReading Comprehension

Similar Papers 제목 키워드 기반

Observational Scaling Laws and the Predictability of Language Model Performance

2024-05-17 · Yangjun Ruan, Chris J. Maddison, Tatsunori Hashimoto

Understanding how language model performance varies with scale is critical to benchmark and algorithm development. Scaling laws are one approach to building this understanding, but the requirement of training models acro…

Language ModelingLanguage Modelling

Limits of Estimating Heterogeneous Treatment Effects: Guidelines for Practical Algorithm Design

2018-07-01 · ICML 2018 7 · Ahmed Alaa, Mihaela Schaar

Estimating heterogeneous treatment effects from observational data is a central problem in many domains. Because counterfactual data is inaccessible, the problem differs fundamentally from supervised learning, and e…

counterfactualGaussian ProcessesSelection bias

CAnDOIT: Causal Discovery with Observational and Interventional Data from Time-Series

2024-10-03 · Luca Castri, Sariah Mghames, Marc Hanheide, Nicola Bellotto

The study of cause-and-effect is of the utmost importance in many branches of science, but also for many practical applications of intelligent systems. In particular, identifying causal relationships in situations that i…

Causal DiscoveryTime Series

Locally- but not Globally-identified SVARs

2025-04-02 · Emanuele Bacchiocchi, Toru Kitagawa

This paper analyzes Structural Vector Autoregressions (SVARs) where identification of structural parameters holds locally but not globally. In this case there exists a set of isolated structural parameter points that are…

Causality-Encoded Diffusion Models for Interventional Sampling and Edge Inference

2026-04-23 · Li Chen, Xiaotong Shen, Wei Pan arxiv

Standard diffusion models are flexible estimators of complex distributions, but they do not encode causal structures and therefore do not by themselves support causal analysis. We propose a causality-encoded diffusion fr…