paper-with-me

홈 › Papers

TTRL: Test-Time Reinforcement Learning

2025-04-22 · Yuxin Zuo, Kaiyan Zhang, Shang Qu, Li Sheng, Xuekai Zhu, Biqing Qi, Youbang Sun, Ganqu Cui, Ning Ding, BoWen Zhou

This paper investigates Reinforcement Learning (RL) on data without explicit labels for reasoning tasks in Large Language Models (LLMs). The core challenge of the problem is reward estimation during inference while not having access to ground-truth information. While this setting appears elusive, we find that common practices in Test-Time Scaling (TTS), such as majority voting, yield surprisingly effective rewards suitable for driving RL training. In this work, we introduce Test-Time Reinforcement Learning (TTRL), a novel method for training LLMs using RL on unlabeled data. TTRL enables self-evolution of LLMs by utilizing the priors in the pre-trained models. Our experiments demonstrate that TTRL consistently improves performance across a variety of tasks and models. Notably, TTRL boosts the pass@1 performance of Qwen-2.5-Math-7B by approximately 159% on the AIME 2024 with only unlabeled test data. Furthermore, although TTRL is only supervised by the Maj@N metric, TTRL has demonstrated performance to consistently surpass the upper limit of the initial model, and approach the performance of models trained directly on test data with ground-truth labels. Our experimental findings validate the general effectiveness of TTRL across various tasks, and highlight TTRL's potential for broader tasks and domains. GitHub: https://github.com/PRIME-RL/TTRL

📄 PDF Abstract BibTeX arXiv:2504.16084

Code (3)

prime-rl/ttrl 공식 구현 pytorch
tsinghuac3i/awesome-rl-reasoning-recipes 공식 구현
qingyangzhang/empo pytorch

Tasks

Mathreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models

2025-11-09 · Peyman Hosseini, Ondrej Bohdal, Taha Ceritli, Ignacio Castro 외 arxiv

Test-time Reinforcement Learning (TTRL) has shown promise in adapting foundation models for complex tasks at test-time, resulting in large performance improvements. TTRL leverages an elegant two-phase sampling strategy: …

Reinforcement Learning

Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models

2026-03-16 · Lit Sin Tan, Junzhe Chen, Xiaolong Fu, Lichen Ma 외 arxiv

Existing test-time scaling (TTS) methods for unified multimodal models (UMMs) in text-to-image (T2I) generation primarily rely on search or sampling strategies that produce only instance-level improvements, limiting the …

Reinforcement Learning

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

2026-08-04 · Kunbin Xu, Xingzuo Li, Xuefeng Bai, Kehai Chen arxiv

Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with pseudo-labels constructed through majority voting. While effective, th…

Reinforcement Learning

MAPLE: Elevating Medical Reasoning from Statistical Consensus to Process-Led Alignment

2026-03-09 · Kailong Fan, Anqi Pu, Yichen Wu, Wanhua Li 외 arxiv

Recent advances in medical large language models have explored Test-Time Reinforcement Learning (TTRL) to enhance reasoning. However, standard TTRL often relies on majority voting (MV) as a heuristic supervision signal, …

Reinforcement Learning

Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning

2026-01-14 · Zhiyuan Hu, Yunhai Hu, Juncheng Liu, Shuyue Stella Li 외 arxiv

Multi-agent systems have evolved into practical LLM-driven collaborators for many applications, gaining robustness from diversity and cross-checking. However, multi-agent RL (MARL) training is resource-intensive and unst…

Reinforcement Learning