paper-with-me

Papers

Few-Shot Action Localization without Knowing Boundaries

2021-06-08 · Ting-Ting Xie, Christos Tzelepis, Fan Fu, Ioannis Patras

Learning to localize actions in long, cluttered, and untrimmed videos is a hard task, that in the literature has typically been addressed assuming the availability of large amounts of annotated training samples for each class -- either in a fully-supervised setting, where action boundaries are known, or in a weakly-supervised setting, where only class labels are known for each video. In this paper, we go a step further and show that it is possible to learn to localize actions in untrimmed videos when a) only one/few trimmed examples of the target action are available at test time, and b) when a large collection of videos with only class label annotation (some trimmed and some weakly annotated untrimmed ones) are available for training; with no overlap between the classes used during training and testing. To do so, we propose a network that learns to estimate Temporal Similarity Matrices (TSMs) that model a fine-grained similarity pattern between pairs of videos (trimmed or untrimmed), and uses them to generate Temporal Class Activation Maps (TCAMs) for seen or unseen classes. The TCAMs serve as temporal attention mechanisms to extract video-level representations of untrimmed videos, and to temporally localize actions at test time. To the best of our knowledge, we are the first to propose a weakly-supervised, one/few-shot action localization network that can be trained in an end-to-end fashion. Experimental results on THUMOS14 and ActivityNet1.2 datasets, show that our method achieves performance comparable or better to state-of-the-art fully-supervised, few-shot learning methods.

📄 PDF Abstract BibTeX arXiv:2106.04150

Code (1)

june01/wfsal-icmr21 공식 구현 pytorch

Tasks

Action LocalizationFew-Shot Learning

Similar Papers 제목 키워드 기반

Localizing the Common Action Among a Few Videos

2020-08-13 · ECCV 2020 8 · Pengwan Yang, Vincent Tao Hu, Pascal Mettes, Cees G. M. Snoek

This paper strives to localize the temporal extent of an action in a long untrimmed video. Where existing work leverages many examples with their start, their ending, and/or the class of the action during training time, …

Action Localization

FMI-TAL: Few-shot Multiple Instances Temporal Action Localization by Probability Distribution Learning and Interval Cluster Refinement

2024-08-25 · Fengshun Wang, Qiurui Wang, Yuting Wang

The present few-shot temporal action localization model can't handle the situation where videos contain multiple action instances. So the purpose of this paper is to achieve manifold action instances localization in a le…

Action LocalizationFew Shot Temporal Action LocalizationTemporal Action Localization

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

2026-08-24 · Xiwei Liu, Yulong Li, Xinlin Zhuang, Xuhui Li 외 arxiv

Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires …

Object LocalizationSpatial Reasoning

Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey

2023-10-19 · Oriane Siméoni, Éloi Zablocki, Spyros Gidaris, Gilles Puy 외

The recent enthusiasm for open-world vision systems show the high interest of the community to perform perception tasks outside of the closed-vocabulary benchmark setups which have been so popular until now. Being able t…

ObjectObject LocalizationUnsupervised Object Localization

Sequential Attention Source Identification Based on Feature Representation

2023-06-28 · Dongpeng Hou, Zhen Wang, Chao GAO, Xuelong Li

Snapshot observation based source localization has been widely studied due to its accessibility and low cost. However, the interaction of users in existing methods does not be addressed in time-varying infection scenario…

DecoderGraph AttentionInductive Learning