paper-with-me

홈 › Papers

ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks

2025-08-03 · Philip Schroeder, Ondrej Biza, Thomas Weng, Hongyin Luo, James Glass arxiv

Vision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their utility in embodied settings, which require reasoning over long frame sequences from a continuous stream of visual input at each moment of a task attempt. To address this limitation, we propose ROVER (Reasoning Over VidEo Recursively), a framework that enables the model to recursively decompose long-horizon video trajectories into segments corresponding to shorter subtasks within the trajectory. In doing so, ROVER facilitates more focused and accurate reasoning over temporally localized frame sequences without losing global context. We evaluate ROVER, implemented using an in-context learning approach, on diverse OpenX Embodiment videos and on a new dataset derived from RoboCasa that consists of 543 videos showing both expert and perturbed non-expert trajectories across 27 robotic manipulation tasks. ROVER outperforms strong baselines across three video reasoning tasks: task progress estimation, frame-level natural language reasoning, and video question answering. We observe that, by reducing the number of frames the model reasons over at each timestep, ROVER mitigates hallucinations, especially during unexpected or non-optimal moments of a trajectory. In addition, by enabling the implementation of a subtask-specific sliding context window, ROVER's time complexity scales linearly with video length, an asymptotic improvement over baselines. Demos, code, and data available at: https://rover-vlm.github.io

📄 PDF Abstract BibTeX arXiv:2508.01943

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

Hilbert: Recursively Building Formal Proofs with Informal Reasoning

2025-09-26 · Sumanth Varambally, Thomas Voice, Yanchao Sun, Zhifeng Chen 외 arxiv

Large Language Models (LLMs) demonstrate impressive mathematical reasoning abilities, but their solutions frequently contain errors that cannot be automatically checked. Formal theorem proving systems such as Lean 4 offe…

Mathematical Reasoning

MerLean-Prover: A Recursive Looping Harness for Lean 4 Theorem Proving

2026-05-26 · Jinzheng Li, Zeru Zhu, Yuanjie Ren arxiv

MerLean-Prover is an end-to-end Lean4 theorem prover that replaces sorry declarations with kernel-checkable proofs. It is built from three agent types (Planning, Check, and Lean) composed by a recursive outer loop whose …

DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition

2025-04-30 · Z. Z. Ren, Zhihong Shao, Junxiao Song, Huajian Xin 외

We introduce DeepSeek-Prover-V2, an open-source large language model designed for formal theorem proving in Lean 4, with initialization data collected through a recursive theorem proving pipeline powered by DeepSeek-V3. …

Automated Theorem ProvingLarge Language ModelMathematical Reasoning

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

2025-07-31 · Luoxin Chen, Jinming Gu, Liankai Huang, Wenhao Huang 외 arxiv

LLMs have demonstrated strong mathematical reasoning abilities by leveraging reinforcement learning with long chain-of-thought, yet they continue to struggle with theorem proving due to the lack of clear supervision sign…

Automated Theorem ProvingReinforcement LearningMathematical Reasoning

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

2026-08-24 · Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu 외 arxiv

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active tem…

Reinforcement LearningVideo Grounding