paper-with-me

홈 › Papers

WIST: Web-Grounded Iterative Self-Play Tree for Domain-Targeted Reasoning Improvement

2026-03-22 · Fangyuan Li, Pengfei Li, Shijie Wang, Junqi Gao, Jianxing Liu, Biqing Qi, Yuqiang Li arxiv

Recent progress in reinforcement learning with verifiable rewards (RLVR) offers a practical path to self-improvement of language models, but existing methods face a key trade-off: endogenous self-play can drift over iterations, while corpus-grounded approaches rely on curated data environments. We present \textbf{WIST}, a \textbf{W}eb-grounded \textbf{I}terative \textbf{S}elf-play \textbf{T}ree framework for domain-targeted reasoning improvement that learns directly from the open web without requiring any pre-arranged domain corpus. WIST incrementally expands a domain tree for exploration, and retrieves and cleans path-consistent web corpus to construct a controllable training environment. It then performs Challenger--Solver self-play with verifiable rewards, and feeds learnability signals back to update node posteriors and guide subsequent exploration through an adaptive curriculum. Across four backbones, WIST consistently improves over the base models and typically outperforms both purely endogenous self-evolution and corpus-grounded self-play baselines, with the Overall gains reaching \textbf{+9.8} (\textit{Qwen3-4B-Base}) and \textbf{+9.7} (\textit{OctoThinker-8B}). WIST is also domain-steerable, improving \textit{Qwen3-8B-Base} by \textbf{+14.79} in medicine and \textit{Qwen3-4B-Base} by \textbf{+5.28} on PhyBench. Ablations further confirm the importance of WIST's key components for stable open-web learning. Our Code is available at https://github.com/lfy-123/WIST.

📄 PDF Abstract BibTeX arXiv:2603.22352

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Improving Constrained Language Generation via Self-Distilled Twisted Sequential Monte Carlo

2025-07-03 · Sooyeon Kim, Giung Nam, Byoungwoo Park, Juho Lee arxiv

Recent work has framed constrained text generation with autoregressive language models as a probabilistic inference problem. Among these, Zhao et al. (2024) introduced a promising approach based on twisted Sequential Mon…

Text Generation

TWIST: Two-Way Inter-Label Self-Training for Semi-Supervised 3D Instance Segmentation

2022-01-01 · CVPR 2022 1 · Ruihang Chu, Xiaoqing Ye, Zhengzhe Liu, Xiao Tan 외

We explore the way to alleviate the label-hungry problem in a semi-supervised setting for 3D instance segmentation. To leverage the unlabeled data to boost model performance, we present a novel Two-Way Inter-label Se…

3D Instance SegmentationDenoisingInstance SegmentationPseudo Label+1

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

2026-07-17 · Ce Zhang, Ziyang Wang, Yulu Pan, Oluwatumininu Oguntola 외 arxiv

Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recent agentic methods frame this task as mul…

Video Question AnsweringReinforcement Learning

Manipulating the Distributions of Experience used for Self-Play Learning in Expert Iteration

2020-05-30 · Dennis J. N. J. Soemers, Éric Piette, Matthew Stephenson, Cameron Browne

Expert Iteration (ExIt) is an effective framework for learning game-playing policies from self-play. ExIt involves training a policy to mimic the search behaviour of a tree search algorithm - such as Monte-Carlo tree sea…

Board Games

TwistList: Resources and Baselines for Tongue Twister Generation

2023-06-06 · Tyler Loakman, Chen Tang, Chenghua Lin

Previous work in phonetically-grounded language generation has mainly focused on domains such as lyrics and poetry. In this paper, we present work on the generation of tongue twisters - a form of language that is require…

Text Generation