paper-with-me

Papers

Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking

2026-06-05 · Fan Zhang, Vireo Zhang, Shengju Qian, Haoxuan Li, Zheng Lian, Hao Wu, Yuan Gao, Xinyu Geng, Xin Wang, Pheng-Ann Heng arxiv

Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with recent efforts shifting from purely text-based information seeking to multimodal settings. However, existing agentic workflows are largely aligned with evidence accumulation models, which linearly aggregate evidence and lack principled mechanisms for handling contradictory information across heterogeneous modalities. Towards this end, we propose Struct-Searcher, a structural agentic workflow grounded in belief revision theory that explicitly maintains an evolving multimodal structural graph throughout the reasoning process, enabling effective conflict-aware multimodal deep information seeking. Extensive experiments across multiple benchmark datasets and backbone models demonstrate that Struct-Searcher is (1) plug-and-play and model-agnostic, yielding an average relative accuracy improvement of 17.2% on BrowseComp-VL across five different backbones. (2) top-performing, consistently outperforming state-of-the-art vision-language models (VLMs) and deep research agents, with relative accuracy improvements of 3.7% on MM-BrowseComp, 1.5% on HLE-VL, and 0.7% on BrowseComp-VL over the second-best competing approach.

📄 PDF Abstract BibTeX arXiv:2606.07689

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Introducing LongCat-Flash-Thinking: A Technical Report

2025-09-23 · Meituan LongCat Team, Anchun Gui, Bei Li, Bingyang Tao 외 arxiv

We present LongCat-Flash-Thinking, an efficient 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model. Its advanced capabilities are cultivated through a meticulously crafted training process, beginn…

Reinforcement Learning

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

2026-05-27 · Minki Kang, Shizhe Diao, Ryo Hachiuma, Sung Ju Hwang 외 arxiv

Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning alone often cannot resolve. Agentic reasoning therefore interleaves …

HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness

2026-05-04 · Jianing Wang, Linsen Guo, Zhengyu Chen, Qi Guo 외 arxiv

Recent advances in agentic harness with orchestration frameworks that coordinate multiple agents with memory, skills, and tool use have achieved remarkable success in complex reasoning tasks. However, the underlying mech…

Reinforcement Learning

DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents

2026-02-03 · Jiahao Zhao, Shaoxuan Xu, Zhongxiang Sun, Fengqi Zhu 외 arxiv

Recently, Diffusion Large Language Models (dLLMs) have demonstrated unique efficiency advantages, enabled by their inherently parallel decoding mechanism and flexible generation paradigm. Meanwhile, despite the rapid adv…

I2I-STRADA -- Information to Insights via Structured Reasoning Agent for Data Analysis

2025-07-23 · SaiBarath Sundar, Pranav Satheesan, Udayaadithya Avadhanam arxiv

Recent advances in agentic systems for data analysis have emphasized automation of insight generation through multi-agent frameworks, and orchestration layers. While these systems effectively manage tasks like query tran…