paper-with-me

Papers

Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering

2021-06-25 · Long Hoang Dang, Thao Minh Le, Vuong Le, Truyen Tran

Video Question Answering (Video QA) is a powerful testbed to develop new AI capabilities. This task necessitates learning to reason about objects, relations, and events across visual and linguistic domains in space-time. High-level reasoning demands lifting from associative visual pattern recognition to symbol-like manipulation over objects, their behavior and interactions. Toward reaching this goal we propose an object-oriented reasoning approach in that video is abstracted as a dynamic stream of interacting objects. At each stage of the video event flow, these objects interact with each other, and their interactions are reasoned about with respect to the query and under the overall context of a video. This mechanism is materialized into a family of general-purpose neural units and their multi-level architecture called Hierarchical Object-oriented Spatio-Temporal Reasoning (HOSTR) networks. This neural model maintains the objects' consistent lifelines in the form of a hierarchically nested spatio-temporal graph. Within this graph, the dynamic interactive object-oriented representations are built up along the video sequence, hierarchically abstracted in a bottom-up manner, and converge toward the key information for the correct answer. The method is evaluated on multiple major Video QA datasets and establishes new state-of-the-arts in these tasks. Analysis into the model's behavior indicates that object-oriented reasoning is a reliable, interpretable and efficient approach to Video QA.

📄 PDF Abstract BibTeX arXiv:2106.13432

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectQuestion AnsweringVideo Question Answering

Similar Papers 제목 키워드 기반

VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation

2026-03-28 · Jihwan Hong, Jaeyoung Do arxiv

Referring Video Object Segmentation (RVOS) aims to segment target objects in videos based on natural language descriptions. However, fixed keyframe-based approaches that couple a vision language model with a separate pro…

Referring Video Object Segmentation

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos

2025-12-03 · Qi'ao Xu, Tianwen Qian, Yuqian Fu, Kailing Li 외 arxiv

A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG). Despite recent progress, existing STV…

Spatio-Temporal Video Grounding

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning

2026-04-12 · Xiaoda Yang, Yuxiang Liu, Shenzhou Gao, Can Wang 외 arxiv

Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for embodied, egocentric tasks. A major source of failure is their relian…

Logical Reasoning

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration

2026-04-23 · Yuehan Zhu, Jingqi Zhao, Jiawen Zhao, Xudong Mao 외 arxiv

Long-form video understanding remains fundamentally challenged by pervasive spatiotemporal redundancy and intricate narrative dependencies that span extended temporal horizons. While recent structured representations com…

Boundary Detection

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

2025-08-29 · Jinzhou Tang, Jusheng zhang, Sidi Liu, Waikit Xiu 외 arxiv

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene u…

Scene UnderstandingQuestion Answering