paper-with-me

Papers

WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning

2025-09-16 · Kuan Li, Zhongwang Zhang, Huifeng Yin, Rui Ye, Yida Zhao, Liwen Zhang, Litu Ou, Dingchu Zhang, Xixi Wu, Jialong Wu, Xinyu Wang, Zile Qiao, Zhen Zhang, Yong Jiang, Pengjun Xie, Fei Huang, Jingren Zhou arxiv

Transcending human cognitive limitations represents a critical frontier in LLM training. Proprietary agentic systems like DeepResearch have demonstrated superhuman capabilities on extremely complex information-seeking benchmarks such as BrowseComp, a feat previously unattainable. We posit that their success hinges on a sophisticated reasoning pattern absent in open-source models: the ability to systematically reduce extreme uncertainty when navigating vast information landscapes. Based on this insight, we introduce WebSailor, a complete post-training methodology designed to instill this crucial capability. Our approach involves generating novel, high-uncertainty tasks through structured sampling and information obfuscation, RFT cold start, and an efficient agentic RL training algorithm, Duplicating Sampling Policy Optimization (DUPO). With this integrated pipeline, WebSailor significantly outperforms all open-source agents in complex information-seeking tasks, matching proprietary agents' performance and closing the capability gap.

📄 PDF Abstract BibTeX arXiv:2509.13305

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

WebSailor: Navigating Super-human Reasoning for Web Agent

2025-07-03 · Kuan Li, Zhongwang Zhang, Huifeng Yin, Liwen Zhang 외

Transcending human cognitive limitations represents a critical frontier in LLM training. Proprietary agentic systems like DeepResearch have demonstrated superhuman capabilities on extremely complex information-seeking be…

Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training

2025-08-01 · Tianqing Fang, Zhisong Zhang, Xiaoyang Wang, Rui Wang 외 arxiv

General AI Agents are increasingly recognized as foundational frameworks for the next generation of artificial intelligence, enabling complex reasoning, web interaction, coding, and autonomous research capabilities. Howe…

CoA: Towards Real Image Dehazing via Compression-and-Adaptation

2025-01-01 · CVPR 2025 1 · Long Ma, Yuxin Feng, Yan Zhang, JinYuan Liu 외

Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world…

Image DehazingModel Compression

Bridging the Semantic Chasm: Synergistic Conceptual Anchoring for Generalized Few-Shot and Zero-Shot OOD Perception

2026-01-30 · Alexandros Christoforos, Sarah Jenkins, Michael Brown, Tuan Pham 외 arxiv

This manuscript presents a pioneering Synergistic Neural Agents Network (SynerNet) framework designed to mitigate the phenomenon of cross-modal alignment degeneration in Vision-Language Models (VLMs) when encountering Ou…

The CHASM-SWPC Dataset for Coronal Hole Detection & Analysis

2025-11-18 · Cutter Beck, Evan Smith, Khagendra Katuwal, Rudra Kafle 외 arxiv

Coronal holes (CHs) are low-activity, low-density solar coronal regions with open magnetic field lines (Cranmer 2009). In the extreme ultraviolet (EUV) spectrum, CHs appear as dark patches. Using daily hand-drawn maps fr…