paper-with-me

Papers

Studying the Korean Word-Chain Game with RLVR: Mitigating Reward Conflicts via Curriculum Learning

2025-10-03 · Donghwan Rho arxiv

Reinforcement learning with verifiable rewards (RLVR) is a promising approach for training large language models (LLMs) with stronger reasoning abilities. It has also been applied to a variety of logic puzzles. In this work, we study the Korean word-chain game using RLVR. We show that rule-derived rewards can naturally conflict, and demonstrate through experiments that a curriculum-learning scheme mitigates these conflicts. Our findings motivate further studies of puzzle tasks in diverse languages.

📄 PDF Abstract BibTeX arXiv:2510.03394

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation

2026-01-26 · Yuxin Jiang, Yufei Wang, Qiyuan Zhang, Xingshan Zeng 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) succeeds in reasoning tasks (e.g., math and code) by checking the final verifiable answer (i.e., a verifiable dot signal). However, extending this paradigm to open-en…

Reinforcement Learning

Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

2025-06-02 · Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 외

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs), while its mechanisms are not yet well understood. In this …

Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning

2025-07-05 · Nayeon Kim, Eojin Jeon, Jun-Hyung Park, SangKeun Lee arxiv

In this study, we introduce KOPL, a novel framework for handling Korean OOV words with Phoneme representation Learning. Our work is based on the linguistic property of Korean as a phonemic script, the high correlation be…

Representation Learning

Studying Semantic Chain Shifts with Word2Vec: FOOD\textgreaterMEAT\textgreaterFLESH

2019-08-01 · WS 2019 8 · Richard Zimmermann

Word2Vec models are used to study the semantic chain shift FOOD{\textgreater}MEAT{\textgreater}FLESH in the history of English, c. 1425-1925. The development stretches out over a long time, starting before 1500, and may …

Recognition Method of Important Words in Korean Text based on Reinforcement Learning

2020-10-01 · CCL 2020 10 · Yang Feiyang, Zhao Yahui, Cui Rongyi

The manual labeling work for constructing the Korean corpus is too time-consuming and laborious. It is difficult for low-minority languages to integrate resources. As a result, the research progress of Korean language in…

Classificationreinforcement-learningReinforcement LearningReinforcement Learning (RL)+4