paper-with-me

Papers

TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes

2025-02-04 · Xingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer, MingYu Liu, Hu Cao, Jiajie Zhang, Venkatnarayanan Lakshminarasimhan, Leah Strand, Alois C. Knoll

We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatio-temporal object expressions, TUMTraffic-VideoQA unifies three essential tasks-multiple-choice video question answering, referred object captioning, and spatio-temporal object grounding-within a cohesive evaluation framework. We further introduce the TUMTraffic-Qwen baseline model, enhanced with visual token sampling strategies, providing valuable insights into the challenges of fine-grained spatio-temporal reasoning. Extensive experiments demonstrate the dataset's complexity, highlight the limitations of existing models, and position TUMTraffic-VideoQA as a robust foundation for advancing research in intelligent transportation systems. The dataset and benchmark are publicly available to facilitate further exploration.

📄 PDF Abstract BibTeX arXiv:2502.02449

Code (1)

TraffiX-VideoQA/TUMTraffic-VideoQA-Baseline

Tasks

Autonomous DrivingMultiple-choiceObjectQuestion AnsweringTraffic Object DetectionVideo Question AnsweringVideo Understanding

Similar Papers 제목 키워드 기반

Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering

2024-04-05 · Lili Liang, Guanglu Sun, Jin Qiu, Lizhong Zhang

Compositional spatio-temporal reasoning poses a significant challenge in the field of video question answering (VideoQA). Existing approaches struggle to establish effective symbolic reasoning structures, which are cruci…

Question AnsweringVideo Question Answering

InterAct-Video: Reasoning-Rich Video QA for Urban Traffic

2025-07-19 · Joseph Raj Vishal, Divesh Basina, Rutuja Patil, Manas Srinivas Gowda 외 arxiv

Traffic monitoring is crucial for urban mobility, road safety, and intelligent transportation systems (ITS). Deep learning has advanced video-based traffic monitoring through video question answering (VideoQA) models, en…

Video Question Answering

UDVideoQA: A Traffic Video Question Answering Dataset for Multi-Object Spatio-Temporal Reasoning in Urban Dynamics

2026-02-24 · Joseph Raj Vishal, Nagasiri Poluri, Katha Naik, Rutuja Patil 외 arxiv

Understanding the complex, multi-agent dynamics of urban traffic remains a fundamental challenge for video language models. This paper introduces Urban Dynamics VideoQA, a benchmark dataset that captures the unscripted r…

Video Question AnsweringMultimodal ReasoningQuestion GenerationVisual Grounding

Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question Answering

2023-07-25 · Yi Cheng, Hehe Fan, Dongyun Lin, Ying Sun 외

The main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing graph-based methods for VideoQA usually …

graph constructionQuestion AnsweringRelationVideo Question Answering

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks

2024-12-02 · Joseph Raj Vishal, Divesh Basina, Aarya Choudhary, Bharatesh Chakravarthi

Recent advances in video question answering (VideoQA) offer promising applications, especially in traffic monitoring, where efficient video interpretation is critical. Within ITS, answering complex, real-time queries lik…

Multi-Object TrackingObject TrackingQuestion AnsweringVideo Question Answering