paper-with-me

홈 › Papers

Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

2025-10-23 · Jiahao Meng, Xiangtai Li, Haochen Wang, Yue Tan, Tao Zhang, Lingdong Kong, Yunhai Tong, Anran Wang, Zhiyang Teng, Yujing Wang, Zhuochen Wang arxiv

Most video reasoning models only generate textual reasoning traces without indicating when and where key evidence appears. Recent models such as OpenAI-o3 have sparked wide interest in evidence-centered reasoning for images, yet extending this ability to videos is more challenging due to the need for joint temporal tracking and spatial localization across dynamic scenes. We introduce Open-o3-Video, a non-agent framework that integrates explicit spatio-temporal evidence into video reasoning by highlighting key timestamps, objects, and bounding boxes, making the reasoning process traceable and verifiable. To enable this capability, we first construct high-quality datasets STGR that provide unified spatio-temporal supervision, which is absent in existing resources. We further adopt a cold-start reinforcement learning strategy with specially designed rewards that jointly encourage answer accuracy, temporal alignment, and spatial precision. On the V-STAR benchmark, Open-o3-Video achieves state-of-the-art performance, improving mAM by 14.4% and mLGM by 24.2% over the Qwen2.5-VL baseline, and shows consistent gains across a range of video understanding benchmarks. Beyond accuracy, the grounded reasoning traces produced by Open-o3-Video support confidence-aware test-time scaling, improving answer reliability.

📄 PDF Abstract BibTeX arXiv:2510.20579

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

2025-08-21 · Pengcheng Fang, Yuxia Chen, Rui Guo arxiv

Understanding videos requires more than answering open ended questions, it demands the ability to pinpoint when events occur and how entities interact across time. While recent Video LLMs have achieved remarkable progres…

Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning

2025-12-05 · Chinthani Sugandhika, Chen Li, Deepu Rajan, Basura Fernando arxiv

Large Video-Language Models (Video-LMs) have achieved impressive progress in multimodal understanding, yet their reasoning remains weakly grounded in space and time. We present Know-Show, a new benchmark designed to eval…

Multimodal Reasoning

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

2026-06-23 · Linpeng Huang, Weixing Chen, Zexin Chen, Yang Liu 외 arxiv

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctne…

Video Question AnsweringAnswer Generation

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living

2026-06-18 · Arkaprava Sinha, Dominick Reilly, Siddharth Krishnan, Hieu Le 외 arxiv

Long Video Question Answering (LVQA) requires identifying sparse, query-relevant evidence within hours-long untrimmed videos. Existing approaches either process videos densely with large vision-language models (VLMs), in…

Video Question Answering

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

2026-06-24 · Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang, Alyssa M. Hardin 외 arxiv

We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly availabl…

Question Answering