paper-with-me

홈 › Papers

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

2026-08-04 · Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao hf

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.

📄 PDF Abstract BibTeX arXiv:2608.03979

Code (3)

BaiShuanghao/my_arXiv_daily ★ 206
Osilly/Vision-DeepResearch ★ 669
Tavish9/awesome-daily-AI-arxiv ★ 113

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Dingtalk DeepResearch: A Unified Multi Agent Framework for Adaptive Intelligence in Enterprise Environments

2025-10-22 · Mengyuan Chen, Chengjun Dai, Xinyang Dong, Chengzhe Feng 외 arxiv

We present Dingtalk DeepResearch, a unified multi agent intelligence framework for real world enterprise environments, delivering deep research, heterogeneous table reasoning, and multimodal report generation.

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework

2025-06-03 · Zhaorui Yang, Bo Pan, Han Wang, Yiyao Wang 외

Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep researc…

Retrieval-augmented Generation

VideoDeepResearch: Long Video Understanding With Agentic Tool Using

2025-06-12 · Huaying Yuan, Zheng Liu, Junjie Zhou, Ji-Rong Wen 외

Long video understanding (LVU) presents a significant challenge for current multi-modal large language models (MLLMs) due to the task's inherent complexity and context window constraint. It is widely assumed that address…

MMEVideo MMEVideo Understanding

Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models

2026-01-29 · Wenxuan Huang, Yu Zeng, Qiuchen Wang, Zhen Fang 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity of their internal world knowledge, prior work has proposed augmenting M…

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

2026-02-03 · Changze Lv, Jie Zhou, Wentao Zhao, Jingwen Xu 외 arxiv

Nowadays, developing reliable DeepResearch-style long-form report generation remains challenging, as training and evaluation lack verifiable reward signals. Accordingly, rubric-based evaluation has become a common practi…

Reinforcement Learning