paper-with-me

홈 › Papers

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

2026-05-21 · Dazhao Du, Jian Liu, Jialong Qin, Tao Han, Bohai Gu, Fangqi Zhu, Yujia Zhang, Eric Liu, Xi Chen, Song Guo arxiv

Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather than by tracking spatiotemporal dynamics. This issue is exacerbated in RL post-training, where correctness-only rewards can further reinforce shortcut policies that obtain high reward without tracking video dynamics. We address this by asking a controlled counterfactual question: if the visual world changed while the question remained fixed, should the answer change or stay the same? Based on this view, we propose \textbf{Counterfactual Relational Policy Optimization (CRPO)}, a dual-branch RL framework for improving \emph{spatiotemporal sensitivity}. CRPO constructs counterfactual videos through horizontal flips and temporal reversals, trains on both original and counterfactual branches, and introduces a \textbf{Counterfactual Relation Reward (CRR)} between their answers. CRR encourages answers to change for dynamic questions and remain unchanged for static questions. This cross-branch constraint makes it difficult for shortcut policies to be consistently rewarded across both branches. To evaluate this property, we introduce \textbf{DyBench}, a paired counterfactual video benchmark with 3,014 videos covering reversible dynamics, moving direction, and event sequence, together with a strict pair-accuracy metric that prevents fixed-answer shortcuts from inflating scores. Experiments show that CRPO outperforms prior RL methods on spatiotemporal-sensitive evaluations while maintaining competitive general video performance. On Qwen3-VL-8B, CRPO improves DyBench P-Acc by +7.7 and TimeBlind I-Acc by +8.2 over the base model, indicating improved spatiotemporal sensitivity rather than stronger reliance on static shortcuts. The project website can be found at https://ddz16.github.io/crpo.github.io/ .

📄 PDF Abstract BibTeX arXiv:2605.21988

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Generative Spatiotemporal Intent Sequence Recommendation via Implicit Reasoning in Amap

2026-05-27 · Sicong Wang, Ruiting Dong, Yue Liu, Bowen Zheng 외 arxiv

Real-world user behavior rarely consists of isolated actions; instead, it often forms intent flows governed by spatiotemporal dependencies. To provide integrated service recommendations, we focus on the task of Generativ…

General Knowledge

Success is in the Details: Evaluate and Enhance Details Sensitivity of Code LLMs through Counterfactuals

2025-05-20 · Xianzhen Luo, Qingfu Zhu, Zhiming Zhang, Mingzheng Xu 외

Code Sensitivity refers to the ability of Code LLMs to recognize and respond to details changes in problem descriptions. While current code benchmarks and instruction data focus on difficulty and diversity, sensitivity i…

counterfactualDiversitySensitivity

STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models

2026-04-03 · Linfeng Fan, Yuan Tian, Ziwei Li, Zhiwu Lu arxiv

Video Large Language Models (Video-LLMs) remain prone to spatiotemporal hallucinations, often generating visually unsupported details or incorrect temporal relations. Existing mitigation methods typically treat hallucina…

Visual Grounding

RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos

2025-10-10 · Zixi Yang, Jiapeng Li, Muxi Diao, Yinuo Jing 외 arxiv

Recently, Multi-modal Large Language Models (MLLMs) have demonstrated significant performance across various video understanding tasks. However, their robustness, particularly when faced with manipulated video content, r…

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering

2026-04-02 · Emad Bahrami, Olga Zatsarynna, Parth Pathak, Sunando Sengupta 외 arxiv

We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for video question answering. While group-based policy optimization methods have…

Video Question AnsweringReinforcement Learning