paper-with-me

홈 › Papers

IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards

2025-08-06 · Xu Guo, Tianyi Liang, Tong Jian, Xiaogui Yang, Ling-I Wu, Chenhui Li, Zhihui Lu, Qipeng Guo, Kai Chen arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) improves instruction following capabilities of large language models (LLMs), but suffers from training inefficiency due to inadequate difficulty assessment. Moreover, RLVR is prone to over-optimization, where LLMs exploit verification shortcuts without aligning to the actual intent of user instructions. We introduce Instruction Following Decorator (IFDecorator}, a framework that wraps RLVR training into a robust and sample-efficient pipeline. It consists of three components: (1) a cooperative-adversarial data flywheel that co-evolves instructions and hybrid verifications, generating progressively more challenging instruction-verification pairs; (2) IntentCheck, a bypass module enforcing intent alignment; and (3) trip wires, a diagnostic mechanism that detects reward hacking via trap instructions, which trigger and capture shortcut exploitation behaviors. Our Qwen2.5-32B-Instruct-IFDecorator achieves 87.43% accuracy on IFEval, outperforming larger proprietary models such as GPT-4o. Additionally, we demonstrate substantial improvements on FollowBench while preserving general capabilities. Our trip wires show significant reductions in reward hacking rates. We will release models, code, and data for future research.

📄 PDF Abstract BibTeX arXiv:2508.04632

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningInstruction Following

Results from the Paper

RankTaskDatasetModelMetrics
#3 Instruction Following IFEval Instruction Inst-level loose-accuracy: 87.43

Similar Papers 제목 키워드 기반

Generalizing Verifiable Instruction Following

2025-07-03 · Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison 외 arxiv

A crucial factor for successful human and AI interaction is the ability of language models or chatbots to follow human instructions precisely. A common feature of instructions are output constraints like ``only answer wi…

Reinforcement LearningInstruction Following

Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following

2026-01-08 · Yirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou 외 arxiv

A central belief in scaling reinforcement learning with verifiable rewards for instruction following (IF) tasks is that, a diverse mixture of verifiable hard and unverifiable soft constraints is essential for generalizin…

Reinforcement LearningInstruction Following

Bridging Offline and Online Reinforcement Learning for LLMs

2025-06-26 · Jack Lanchantin, Angelica Chen, Janice Lan, Xian Li 외

We investigate the effectiveness of reinforcement learning methods for finetuning large language models when transitioning from offline to semi-online to fully online regimes for both verifiable and non-verifiable tasks.…

Instruction FollowingMathreinforcement-learningReinforcement Learning

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following

2025-06-11 · Hao Peng, Yunjia Qi, Xiaozhi Wang, Bin Xu 외

Reinforcement learning with verifiable rewards (RLVR) has become a key technique for enhancing large language models (LLMs), with verification engineering playing a central role. However, best practices for RL in instruc…

Instruction Followingreinforcement-learningReinforcement Learning

ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following

2026-02-04 · Yuancheng Yang, Lin Yang, Xu Wang, Chao Tong 외 arxiv

As applications of large language models (LLMs) become increasingly complex, the demand for robust complex instruction following capabilities is growing accordingly. We argue that a thorough understanding of the instruct…

Reinforcement LearningInstruction Following