paper-with-me

Papers

Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question Answering

2023-07-25 · Yi Cheng, Hehe Fan, Dongyun Lin, Ying Sun, Mohan Kankanhalli, Joo-Hwee Lim

The main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing graph-based methods for VideoQA usually ignore keywords in questions and employ a simple graph to aggregate features without considering relative relations between objects, which may lead to inferior performance. In this paper, we propose a Keyword-aware Relative Spatio-Temporal (KRST) graph network for VideoQA. First, to make question features aware of keywords, we employ an attention mechanism to assign high weights to keywords during question encoding. The keyword-aware question features are then used to guide video graph construction. Second, because relations are relative, we integrate the relative relation modeling to better capture the spatio-temporal dynamics among object nodes. Moreover, we disentangle the spatio-temporal reasoning into an object-level spatial graph and a frame-level temporal graph, which reduces the impact of spatial and temporal relation reasoning on each other. Extensive experiments on the TGIF-QA, MSVD-QA and MSRVTT-QA datasets demonstrate the superiority of our KRST over multiple state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2307.13250

Code (0)

등록된 구현이 없습니다.

Tasks

graph constructionQuestion AnsweringRelationVideo Question Answering

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Competition-Aware CPC Forecasting with Near-Market Coverage

2026-03-13 · Sebastian Frey, Edoardo Beccari, Maximilian Kranz, Nicolò Alberto Pellizzari 외 arxiv

Cost-per-click (CPC) in paid search is an auction-generated outcome shaped by a competitive landscape that is only partially observable from any single advertiser's history. From 1.66 billion Google Ads log records for a…

iMOVE: Instance-Motion-Aware Video Understanding

2025-02-17 · Jiaze Li, Yaya Shi, Zongyang Ma, Haoran Xu 외

Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understanding. However, current models struggle t…

Computational EfficiencyVideo Understanding

NETR-Tree: An Eifficient Framework for Social-Based Time-Aware Spatial Keyword Query

2019-08-26 · Xiuqi Huang, Yuanning Gao, Xiaofeng Gao, Guihai Chen

The development of global positioning system stimulates the popularity of location-based social network (LBSN) services. With a large volume of data containing locations, texts, check-in information, and social relations…

Network Embedding

Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation

2025-03-19 · Haoyu Ji, Bowen Chen, Weihong Ren, Wenze Huang 외

Skeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling …

Contrastive LearningSkeleton Based Action SegmentationTAG

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts

2026-07-07 · Haida Feng, Hao Wei, Haolin Wang, Shiwei Li 외 arxiv

Recent Multimodal Large Language Models (MLLMs) struggle to bridge the representational gap between 2D semantic understanding and 3D spatial geometry. Existing 3D-aware models either rely on costly 3D-specific data or ut…

Spatial Reasoning