Studying the Korean Word-Chain Game with RLVR: Mitigating Reward Conflicts via Curriculum Learning
Reinforcement learning with verifiable rewards (RLVR) is a promising approach for training large language models (LLMs) with stronger reasoning abilities. It has also been applied to a variety of logic puzzles. In this work, we study the Korean word-chain game using RLVR. We show that rule-derived rewards can naturally conflict, and demonstrate through experiments that a curriculum-learning scheme mitigates these conflicts. Our findings motivate further studies of puzzle tasks in diverse languages.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningSimilar Papers 제목 키워드 기반
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
Reinforcement learning with verifiable rewards (RLVR) succeeds in reasoning tasks (e.g., math and code) by checking the final verifiable answer (i.e., a verifiable dot signal). However, extending this paradigm to open-en…
Reinforcement LearningBeyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs), while its mechanisms are not yet well understood. In this …
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
In this study, we introduce KOPL, a novel framework for handling Korean OOV words with Phoneme representation Learning. Our work is based on the linguistic property of Korean as a phonemic script, the high correlation be…
Representation LearningStudying Semantic Chain Shifts with Word2Vec: FOOD\textgreaterMEAT\textgreaterFLESH
Word2Vec models are used to study the semantic chain shift FOOD{\textgreater}MEAT{\textgreater}FLESH in the history of English, c. 1425-1925. The development stretches out over a long time, starting before 1500, and may …
Recognition Method of Important Words in Korean Text based on Reinforcement Learning
The manual labeling work for constructing the Korean corpus is too time-consuming and laborious. It is difficult for low-minority languages to integrate resources. As a result, the research progress of Korean language in…
Classificationreinforcement-learningReinforcement LearningReinforcement Learning (RL)+4