paper-with-me

홈 › Papers

Enhancing Japanese Large Language Models with Reasoning Vectors

2025-08-04 · Carolina Minami Oguchi, Leo Wei, Koyo Kobayashi, Hsin-Tai Wu, Dipak Ghosal arxiv

Post-training methods have improved the performance and enhanced the reasoning capability for mainstream large language models (LLMs), but the same is challenging for Japanese LLMs to achieve due to the amount of resources required. Inspired by task vectors that extract the change of weights before and after training, specifically for a certain task, we obtain reasoning vectors from reasoning LLMs and apply them to Japanese LLMs to boost their performance. While the resources available present a challenge to improve Japanese LLMs, we present a simple and effective way to obtain high improvement and hope to inspire for other languages.

📄 PDF Abstract BibTeX arXiv:2508.02913

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory

2026-04-01 · Taihei Shiotani, Masahiro Kaneko, Naoaki Okazaki arxiv

In enhancing the fairness of Large Language Models (LLMs), evaluating social biases rooted in the cultural contexts of specific linguistic regions is essential. However, most existing Japanese benchmarks heavily rely on …

Cost of Reasoning in non-English Languages: A Case Study on Japanese

2026-07-11 · Yuu Jinnai arxiv

Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training data is most abundant. However, reasoning trace is a clue for model int…

LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment

2025-03-24 · Jong Myoung Kim, Young-Jun Lee, Ho-Jin Choi, SangKeun Jung

While Large Language Models have gained attention, many service developers still rely on embedding-based models due to practical constraints. In such cases, the quality of fine-tuning data directly impacts performance, a…

Language ModelingLanguage Modelling

Why We Build Local Large Language Models: An Observational Analysis from 35 Japanese and Multilingual LLMs

2024-12-19 · Koshiro Saito, Sakae Mizuki, Masanari Ohi, Taishi Nakamura 외

Why do we build local large language models (LLMs)? What should a local LLM learn from the target language? Which abilities can be transferred from other languages? Do language-specific scaling laws exist? To explore the…

Arithmetic ReasoningCode GenerationQuestion AnsweringReading Comprehension

Analyzing Social Biases in Japanese Large Language Models

2024-06-04 · Hitomi Yanaka, Namgi Han, Ryoma Kumon, Jie Lu 외

With the development of Large Language Models (LLMs), social biases in the LLMs have become a crucial issue. While various benchmarks for social biases have been provided across languages, the extent to which Japanese LL…

Question Answering