paper-with-me

홈 › Papers

Making Qwen3 Think in Korean with Reinforcement Learning

2025-08-14 · Jungyup Lee, Jemin Kim, Sang Park, SeungJae Lee arxiv

We present a two-stage fine-tuning approach to make the large language model Qwen3 14B "think" natively in Korean. In the first stage, supervised fine-tuning (SFT) on a high-quality Korean reasoning dataset establishes a strong foundation in Korean logical reasoning, yielding notable improvements in Korean-language tasks and even some gains in general reasoning ability. In the second stage, we employ reinforcement learning with a customized Group Relative Policy Optimization (GRPO) algorithm to further enhance both Korean reasoning alignment and overall problem-solving performance. We address critical stability challenges in GRPO training - such as reward hacking and policy collapse - by introducing an oracle judge model that calibrates the reward signal. Our approach achieves stable learning (avoiding the collapse observed in naive GRPO) and leads to steady, incremental performance gains. The final RL-tuned model demonstrates substantially improved results on advanced reasoning benchmarks (particularly math and coding tasks) while maintaining knowledge and language proficiency, successfully conducting its internal chain-of-thought entirely in Korean.

📄 PDF Abstract BibTeX arXiv:2508.10355

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningLogical Reasoning

Similar Papers 제목 키워드 기반

On-Policy Delta Distillation for Multilingual Math Reasoning

2026-08-06 · Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han hf

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced var…

Mathematical ReasoningReinforcement Learning

KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination

2026-02-14 · Byungjin Choi, Seongsu Bae, Sunjun Kweon, Edward Choi arxiv

We introduce KorMedMCQA-V, a Korean medical licensing-exam-style multimodal multiple-choice question answering benchmark for evaluating vision-language models (VLMs). The dataset consists of 1,534 questions with 2,043 as…

Question Answering

Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

2026-06-30 · Utsav Garg, Sungjin Hong, Jason Jung, Justin Lee 외 arxiv

We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterprise agents under practical memory and serving constraints. The mode…

Reinforcement LearningMathematical Reasoning

BlueLM-2.5-3B Technical Report

2025-07-08 · Baojiao Xiong, Boheng Chen, Chengzhi Wang, Daxiong Luo 외

We present BlueLM-2.5-3B, a compact and unified dense Multimodal Large Language Model (MLLM) designed for efficient edge-device deployment, offering strong general-purpose and reasoning capabilities. To the best of our k…

Large Language ModelMultimodal Large Language Model

HyperCLOVA X 32B Think

2026-01-03 · NAVER Cloud HyperCLOVA X Team arxiv

In this report, we present HyperCLOVA X 32B Think, a vision-language model designed with particular emphasis on reasoning within the Korean linguistic and cultural context, as well as agentic ability. HyperCLOVA X 32B Th…