paper-with-me

Papers

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

2026-06-03 · Lin Fu, Zheyuan Yang, Yang Wang, Tingyu Song, Arman Cohan, Yilun Zhao arxiv

We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding. It comprises 315K video reasoning examples over 145K newly collected, CC-licensed, expert-domain videos. We develop a human-in-the-loop, skill-oriented example generation pipeline that targets progressively deeper video reasoning capabilities while ensuring the difficulty, diversity, and reliability of both the examples and their CoT rationales. We also curate VideoKR-Eval, a new expert-annotated benchmark where questions require genuine video understanding and knowledge-intensive reasoning rather than textual shortcuts. Our experiments show that, under a standard SFT$\rightarrow$GRPO pipeline, models post-trained on VideoKR outperform prior post-training approaches on knowledge-intensive video reasoning while remaining competitive on general video reasoning, highlighting data design as a key driver of progress in video reasoning. We further conduct comprehensive ablations to isolate the contributions of VideoKR, providing actionable insights for future work.

📄 PDF Abstract BibTeX arXiv:2606.05259

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

2025-02-17 · Guangzhi Sun, Yudong Yang, Jimin Zhuang, Changli Tang 외

While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems a…

Language ModelingLanguage ModellingLarge Language ModelVideo Understanding

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

2025-01-21 · CVPR 2025 1 · Yilun Zhao, Lujing Xie, Haowei Zhang, Guo Gan 외

We introduce MMVU, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. MMVU includes 3,000 expert-annotated questions spanning 27 subjects across four core di…

Video Understanding

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

2026-06-05 · Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao 외 arxiv

Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios. These scenarios require model…

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

2025-12-12 · Han Lin, Xichen Pan, Ziqi Huang, Ji Hou 외 arxiv

Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are…

Text-to-Image GenerationVideo Generation

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning

2026-07-03 · Zhenkun Gao, Yicheng Bao, Jinlong Peng, Xueheng Li 외 arxiv

Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (VDR). However, existing multimodal search agents primarily target stati…

Reinforcement Learning