paper-with-me

홈 › Papers

TTVS: Boosting Self-Exploring Reinforcement Learning via Test-time Variational Synthesis

2026-04-09 · Sikai Bai, Haoxi Li, Jie Zhang, Yongjiang Liu, Song Guo arxiv

Despite significant advances in Large Reasoning Models (LRMs) driven by reinforcement learning with verifiable rewards (RLVR), this paradigm is fundamentally limited in specialized or novel domains where such supervision is prohibitively expensive or unavailable, posing a key challenge for test-time adaptation. While existing test-time methods offer a potential solution, they are constrained by learning from static query sets, risking overfitting to textual patterns. To address this gap, we introduce Test-Time Variational Synthesis (TTVS), a novel framework that enables LRMs to self-evolve by dynamically augmenting the training stream from unlabeled test queries. TTVS comprises two synergistic modules: (1) Online Variational Synthesis, which transforms static test queries into a dynamic stream of diverse, semantically-equivalent variations, enforcing the model to learn underlying problem logic rather than superficial patterns; (2) Test-time Hybrid Exploration, which balances accuracy-driven exploitation with consistency-driven exploration across synthetic variants. Extensive experiments show TTVS yields superior performance across eight model architectures. Notably, using only unlabeled test-time data, TTVS not only surpasses other test-time adaptation methods but also outperforms state-of-the-art supervised RL-based techniques trained on vast, high-quality labeled data.

📄 PDF Abstract BibTeX arXiv:2604.08468

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningTest-time Adaptation

Similar Papers 제목 키워드 기반

Learning Trajectory-Aware Transformer for Video Super-Resolution

2022-04-08 · CVPR 2022 1 · Chengxu Liu, Huan Yang, Jianlong Fu, Xueming Qian

Video super-resolution (VSR) aims to restore a sequence of high-resolution (HR) frames from their low-resolution (LR) counterparts. Although some progress has been made, there are grand challenges to effectively utilize …

Super-ResolutionVideo derainingVideo Super-Resolution

Alleviating the transit timing variation bias in transit surveys. I. RIVERS: Method and detection of a pair of resonant super-Earths around Kepler-1705

2021-11-12 · A. Leleu, G. Chatel, S. Udry, Y. Alibert 외

Transit timing variations (TTVs) can provide useful information for systems observed by transit, as they allow us to put constraints on the masses and eccentricities of the observed planets, or even to constrain the exis…

TAPESTRY: From Geometry to Appearance via Consistent Turntable Videos

2026-03-18 · Yan Zeng, Haoran Jiang, Kaixin Yao, Qixuan Zhang 외 arxiv

Automatically generating photorealistic and self-consistent appearances for untextured 3D models is a critical challenge in digital content creation. The advancement of large-scale video generation models offers a natura…

3D ReconstructionVideo Generation

Single Transit Detection In Kepler With Machine Learning And Onboard Spacecraft Diagnostics

2024-03-06 · Matthew T. Hansen, Jason A. Dittmann

Exoplanet discovery at long orbital periods requires reliably detecting individual transits without additional information about the system. Techniques like phase-folding of light curves and periodogram analysis of radia…

Algorithms in Multi-Agent Systems: A Holistic Perspective from Reinforcement Learning and Game Theory

2020-01-17 · Yunlong Lu, Kai Yan

Deep reinforcement learning (RL) has achieved outstanding results in recent years, which has led a dramatic increase in the number of methods and applications. Recent works are exploring learning beyond single-agent scen…

counterfactualDeep Reinforcement LearningMulti-agent Reinforcement Learningreinforcement-learning+2