paper-with-me

홈 › Papers

SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events

2021-03-29 · CVPR 2021 1 · Li Xu, He Huang, Jun Liu

Traffic event cognition and reasoning in videos is an important task that has a wide range of applications in intelligent transportation, assisted driving, and autonomous vehicles. In this paper, we create a novel dataset, SUTD-TrafficQA (Traffic Question Answering), which takes the form of video QA based on the collected 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive capability of causal inference and event understanding models in complex traffic scenarios. Specifically, we propose 6 challenging reasoning tasks corresponding to various traffic scenarios, so as to evaluate the reasoning capability over different kinds of complex yet practical traffic events. Moreover, we propose Eclipse, a novel Efficient glimpse network via dynamic inference, in order to achieve computation-efficient and reliable video reasoning. The experiments show that our method achieves superior performance while reducing the computation cost significantly. The project page: https://github.com/SUTDCV/SUTD-TrafficQA.

📄 PDF Abstract BibTeX arXiv:2103.15538

Code (3)

SUTDCV/SUTD-TrafficQA 공식 구현 pytorch
MarkHershey/arxiv-dl
saccharomycetes/text-based-traffic-understanding pytorch

Tasks

Autonomous VehiclesBenchmarkingCausal InferenceQuestion AnsweringVideo Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Causal inference Causal inference is the process of drawing a conclusion about a causal connection based on the conditions of the occurrence of an effect. The main difference between causal…

Similar Papers 제목 키워드 기반

The Multi-Modal Video Reasoning and Analyzing Competition

2021-08-18 · Haoran Peng, He Huang, Li Xu, Tianjiao Li 외

In this paper, we introduce the Multi-Modal Video Reasoning and Analyzing Competition (MMVRAC) workshop in conjunction with ICCV 2021. This competition is composed of four different tracks, namely, video question answeri…

Action RecognitionPerson Re-IdentificationQuestion AnsweringSkeleton Based Action Recognition+1

Traffic-Domain Video Question Answering with Automatic Captioning

2023-07-18 · Ehsan Qasemi, Jonathan M. Francis, Alessandro Oltramari

Video Question Answering (VidQA) exhibits remarkable potential in facilitating advanced machine reasoning capabilities within the domains of Intelligent Traffic Monitoring and Intelligent Transportation Systems. Neverthe…

Question AnsweringVideo Question Answering

FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering

2025-07-17 · Ju-Young Oh, Ho-Joong Kim, Seong-Whan Lee arxiv

Video question answering (VQA) is a multimodal task that requires the interpretation of a video to answer a given question. Existing VQA methods primarily utilize question and answer (Q&A) pairs to learn the spatio-tempo…

Video Question AnsweringQuestion Generation

Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach

2025-11-18 · Ju-Young Oh arxiv

Conventional VQA approaches primarily rely on question-answer (Q&A) pairs to learn the spatio-temporal dynamics of video content. However, most existing annotations are event-centric, which restricts the model's ability …

Video Question AnsweringQuestion Generation

Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning

2025-09-14 · Waikit Xiu, Qiang Lu, Bingchen Liu, Chen Sun 외 arxiv

For safe and robust autonomous driving, decision-making systems must effectively leverage past experiences to handle the inherent long-tail of traffic scenarios. Case-Based Reasoning (CBR) provides a natural paradigm for…

Autonomous Driving