paper-with-me

홈 › Papers

Tubelets: Unsupervised action proposals from spatiotemporal super-voxels

2016-07-07 · Mihir Jain, Jan van Gemert, Hervé Jégou, Patrick Bouthemy, Cees G. M. Snoek

This paper considers the problem of localizing actions in videos as a sequences of bounding boxes. The objective is to generate action proposals that are likely to include the action of interest, ideally achieving high recall with few proposals. Our contributions are threefold. First, inspired by selective search for object proposals, we introduce an approach to generate action proposals from spatiotemporal super-voxels in an unsupervised manner, we call them Tubelets. Second, along with the static features from individual frames our approach advantageously exploits motion. We introduce independent motion evidence as a feature to characterize how the action deviates from the background and explicitly incorporate such motion information in various stages of the proposal generation. Finally, we introduce spatiotemporal refinement of Tubelets, for more precise localization of actions, and pruning to keep the number of Tubelets limited. We demonstrate the suitability of our approach by extensive experiments for action proposal quality and action localization on three public datasets: UCF Sports, MSR-II and UCF101. For action proposal quality, our unsupervised proposals beat all other existing approaches on the three datasets. For action localization, we show top performance on both the trimmed videos of UCF Sports and UCF101 as well as the untrimmed videos of MSR-II.

📄 PDF Abstract BibTeX arXiv:1607.02003

Code (0)

등록된 구현이 없습니다.

Tasks

Action Localization

Similar Papers 제목 키워드 기반

Object Detection in Videos with Tubelet Proposal Networks

2017-02-21 · CVPR 2017 7 · Kai Kang, Hongsheng Li, Tong Xiao, Wanli Ouyang 외

Object detection in videos has drawn increasing attention recently with the introduction of the large-scale ImageNet VID dataset. Different from object detection in static images, temporal information in videos is vital …

Objectobject-detectionObject DetectionObject Tracking

End-to-End Spatio-Temporal Action Localisation with Video Transformers

2023-04-24 · CVPR 2024 1 · Alexey Gritsenko, Xuehan Xiong, Josip Djolonga, Mostafa Dehghani 외

The most performant spatio-temporal action localisation models use external person proposals and complex external memory banks. We propose a fully end-to-end, purely-transformer based model that directly ingests an input…

Action DetectionAction RecognitionSpatio-Temporal Action Localization

Social Fabric: Tubelet Compositions for Video Relation Detection

2021-08-18 · ICCV 2021 10 · Shuo Chen, Zenglin Shi, Pascal Mettes, Cees G. M. Snoek

This paper strives to classify and detect the relationship between object tubelets appearing within a video as a <subject-predicate-object> triplet. Where existing works treat object proposals or tubelets as single entit…

ObjectRelationTripletVideo Visual Relation Detection+1

From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition

2026-03-27 · Nazia Aslam, Abhisek Ray, Joakim Bruslund Haurum, Lukas Esterle 외 arxiv

Recent advances in large-scale video models have significantly improved video understanding across domains such as surveillance, healthcare, and entertainment. However, these models also amplify privacy risks by encoding…

Action Recognition

Spatiotemporal Learning with Context-aware Video Tubelets for Ultrasound Video Analysis

2025-03-21 · Gary Y. Li, Li Chen, Bryson Hicks, Nikolai Schnittke 외

Computer-aided pathology detection algorithms for video-based imaging modalities must accurately interpret complex spatiotemporal information by integrating findings across multiple frames. Current state-of-the-art metho…

object-detectionObject DetectionVideo Classification