paper-with-me

Papers

STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding

2025-01-01 · CVPR 2025 1 · Aaryan Garg, Akash Kumar, Yogesh S Rawat

In this work, we study Weakly Supervised Spatio-Temporal Video Grounding (WSTVG), a challenging task of localizing subjects spatio-temporally in videos using only textual queries and no bounding box supervision. Inspired by recent advances in vision-language foundation models, we investigate their utility for WSTVG, leveraging their zero-shot grounding capabilities. However, we find that a simple adaptation lacks essential spatio-temporal grounding abilities. To bridge this gap, we introduce Tubelet Referral Grounding (TRG), which connects textual queries to tubelets to enable spatio-temporal predictions. Despite its promise, TRG struggles with compositional action understanding and dense scene scenarios. To address these limitations, we propose STPro, a progressive learning framework with two key modules: Sub-Action Temporal Curriculum Learning (SA-TCL), which incrementally builds compositional action understanding, and Congestion-Guided Spatial Curriculum Learning (CG-SCL), which adapts the model to complex scenes by spatially increasing task difficulty. STPro achieves state-of-the-art results on three benchmark datasets, with improvements of 1.0% on VidSTG-Declarative and 3.0% on HCSTVG-v1.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action UnderstandingSpatio-Temporal Video GroundingVideo Grounding

Similar Papers 제목 키워드 기반

Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts

2024-08-12 · Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang 외

Current weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global …

Anomaly DetectionEvent Detectionobject-detectionObject Detection+2

STIPP: Space-time in situ postprocessing over the French Alps using proper scoring rules

2026-01-06 · David Landry, Isabelle Gouttevin, Hugo Merizen, Claire Monteleoni 외 arxiv

We propose Space-time in situ postprocessing (STIPP), a machine learning model that generates spatio-temporally consistent weather forecasts for a network of station locations. Gridded forecasts from classical numerical …

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

2021-08-09 · Jie Wu, Wei zhang, Guanbin Li, Wenhao Wu 외

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video, WSSTAD aims to localize a spatio-tempor…

Anomaly Detection

UniPose: Unified Human Pose Estimation in Single Images and Videos

2020-01-22 · CVPR 2020 6 · Bruno Artacho, Andreas Savakis

We propose UniPose, a unified framework for human pose estimation, based on our "Waterfall" Atrous Spatial Pooling architecture, that achieves state-of-art-results on several pose estimation metrics. Current pose estimat…

Pose EstimationSkeleton Based Action Recognition

Learning Where and When: Patch-Based Spatiotemporal Localization in Weakly Supervised Video Anomaly Detection

2026-06-28 · Hamza Karim, Nghia Nguyen, Lokman Bekit, Yasin Yilmaz arxiv

Weakly supervised video anomaly detection (WSVAD) has predominantly focused on temporal localization, identifying when anomalies occur while largely neglecting their spatial extent within frames. Yet, spatial localizatio…

Multiple Instance LearningVideo Anomaly Detection