paper-with-me

홈 › Papers

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding

2026-04-09 · Xuezhen Tu, Jingyu Wu, Fangyu Kang, Qingpeng Nong, Kaijin Zhang, Chaoyue Niu, Fan Wu arxiv

Spatio-Temporal Video Grounding requires jointly localizing target objects across both temporal and spatial dimensions based on natural language queries, posing fundamental challenges for existing Multimodal Large Language Models (MLLMs). We identify two core challenges: \textit{entangled spatio-temporal alignment}, arising from coupling two heterogeneous sub-tasks within the same autoregressive output space, and \textit{dual-domain visual token redundancy}, where target objects exhibit simultaneous temporal and spatial sparsity, rendering the overwhelming majority of visual tokens irrelevant to the grounding query. To address these, we propose \textbf{Bridge-STG}, an end-to-end framework that decouples temporal and spatial localization while maintaining semantic coherence. While decoupling is the natural solution to this entanglement, it risks creating a semantic gap between the temporal MLLM and the spatial decoder. Bridge-STG resolves this through two pivotal designs: the \textbf{Spatio-Temporal Semantic Bridging (STSB)} mechanism with Explicit Temporal Alignment (ETA) distills the MLLM's temporal reasoning context into enriched bridging queries as a robust semantic interface; and the \textbf{Query-Guided Spatial Localization (QGSL)} module leverages these queries to drive a purpose-built spatial decoder with multi-layer interactive queries and positive/negative frame sampling, jointly eliminating dual-domain visual token redundancy. Extensive experiments across multiple benchmarks demonstrate that Bridge-STG achieves state-of-the-art performance among MLLM-based methods. Bridge-STG improves average m\_vIoU from $26.4$ to $34.3$ on VidSTG and demonstrates strong cross-task transfer across various fine-grained video understanding tasks under a unified multi-task training regime.

📄 PDF Abstract BibTeX arXiv:2604.08014

Code (0)

등록된 구현이 없습니다.

Tasks

Spatio-Temporal Video GroundingNatural Language Queries

Similar Papers 제목 키워드 기반

AutoSTF: Decoupled Neural Architecture Search for Cost-Effective Automated Spatio-Temporal Forecasting

2024-09-25 · Tengfei Lyu, Weijia Zhang, Jinliang Deng, Hao liu

Spatio-temporal forecasting is a critical component of various smart city applications, such as transportation optimization, energy management, and socio-economic analysis. Recently, several automated spatio-temporal for…

energy managementModel OptimizationNeural Architecture SearchSpatio-Temporal Forecasting

Spatio-Temporal Gating-Adjacency GCN for Human Motion Prediction

2022-03-03 · CVPR 2022 1 · Chongyang Zhong, Lei Hu, Zihao Zhang, Yongjing Ye 외

Predicting future motion based on historical motion sequence is a fundamental problem in computer vision, and it has wide applications in autonomous driving and robotics. Some recent works have shown that Graph Convoluti…

Autonomous DrivingHuman motion predictionmotion predictionPrediction

Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion Recognition

2021-12-16 · CVPR 2022 1 · Benjia Zhou, Pichao Wang, Jun Wan, Yanyan Liang 외

Decoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising perfor…

Hand Gesture Recognition

Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting

2023-12-01 · Haotian Gao, Renhe Jiang, Zheng Dong, Jinliang Deng 외

Spatiotemporal forecasting techniques are significant for various domains such as transportation, energy, and weather. Accurate prediction of spatiotemporal series remains challenging due to the complex spatiotemporal he…

Time SeriesTraffic Prediction

A Decoupled Spatio-Temporal Framework for Skeleton-based Action Segmentation

2023-12-10 · Yunheng Li, Zhongyu Li, ShangHua Gao, Qilong Wang 외

Effectively modeling discriminative spatio-temporal information is essential for segmenting activities in long action sequences. However, we observe that existing methods are limited in weak spatio-temporal modeling capa…

Action SegmentationSkeleton Based Action Segmentation