paper-with-me

홈 › Papers

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering

2026-06-23 · Zhentao Guo, Chen Duan, Tongkun Guan, Zining Wang, Kai Zhou, Pengfei Yan arxiv

Despite remarkable progress in multimodal understanding, current MLLMs still exhibit limitations in video text understanding, particularly when semantics emerge through the integration of temporally distributed textual cues across multiple frames. This perception challenge fundamentally differs from static image text understanding, yet existing datasets fail to capture: the vast majority of questions remain answerable from single frames, inadequately reflecting real-world video text comprehension demands. To address this, we present ViTexQA, a large-scale video-text QA dataset, and FrameThinker for robust multi-frame temporal reasoning. We build ViTexQA via a quality-controlled Chain-of-Thought (CoT) annotation pipeline boosted with temporal constraints; all its QA pairs demand cross-frame text fusion to solve, enforcing true temporal reliance. FrameThinker adopts two-stage training for explicit temporal modeling: CoT-Guided Supervised Fine-Tuning (SFT) generates frame-aware reasoning chains, followed by Temporally-grounded Reinforcement Learning (RL) optimized with multi-frame coherence rewards. Evaluations show our method outperforms SOTA baselines on ViTexQA, lifting ROUGE-L by 6.3%.

📄 PDF Abstract BibTeX arXiv:2606.24602

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningQuestion Answering

Similar Papers 제목 키워드 기반

V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction

2024-12-02 · Zewei Zhou, Hao Xiang, Zhaoliang Zheng, Seth Z. Zhao 외

Vehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception,…

Prediction

CoDynTrust: Robust Asynchronous Collaborative Perception via Dynamic Feature Trust Modulus

2025-02-12 · Yunjiang Xu, Lingzhi Li, Jin Wang, Benyuan Yang 외

Collaborative perception, fusing information from multiple agents, can extend perception range so as to improve perception performance. However, temporal asynchrony in real-world environments, caused by communication del…

Towards All-Day Perception for Off-Road Driving: A Large-Scale Multispectral Dataset and Comprehensive Benchmark

2026-04-30 · Shuo Wang, Jilin Mei, Wenfei Guan, Shuai Wang 외 arxiv

Off-road nighttime autonomous driving suffers from unreliable visible-light perception, making infrared modality crucial for accurate freespace detection. However, progress remains limited due to the scarcity of annotate…

Autonomous Driving

Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception

2023-07-26 · ICCV 2023 1 · Kun Yang, Dingkang Yang, Jingyu Zhang, Mingcheng Li 외

Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However,…

3D Object DetectionAutonomous Vehiclesobject-detectionObject Detection

HENet: Hybrid Encoding for End-to-end Multi-task 3D Perception from Multi-view Cameras

2024-04-03 · Zhongyu Xia, Zhiwei Lin, Xinhao Wang, Yongtao Wang 외

Three-dimensional perception from multi-view cameras is a crucial component in autonomous driving systems, which involves multiple tasks like 3D object detection and bird's-eye-view (BEV) semantic segmentation. To improv…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection+1