paper-with-me

홈 › Papers

Realigning Confidence with Temporal Saliency Information for Point-Level Weakly-Supervised Temporal Action Localization

2024-01-01 · CVPR 2024 1 · Ziying Xia, Jian Cheng, Siyu Liu, Yongxiang Hu, Shiguang Wang, Yijie Zhang, Liwan Dang

Point-level weakly-supervised temporal action localization (P-TAL) aims to localize action instances in untrimmed videos through the use of single-point annotations in each instance. Existing methods predict the class activation sequences without any boundary information and the unreliable sequences result in a significant misalignment between the quality of proposals and their corresponding confidence. In this paper we surprisingly observe the most salient frame tend to appear in the central region of the each instance and is easily annotated by humans. Guided by the temporal saliency information we present a novel proposal-level plug-in framework to relearn the aligned confidence of proposals generated by the base locators. The proposed approach consists of Center Score Learning (CSL) and Alignment-based Boundary Adaptation (ABA). In CSL we design a novel center label generated by the point annotations for predicting aligned center scores. During inference we first fuse the center scores with the predicted action probabilities to obtain the aligned confidence. ABA utilizes the both aligned confidence and IoU information to enhance localization completeness. Extensive experiments demonstrate the generalization and effectiveness of the proposed framework showcasing state-of-the-art or competitive performances across three benchmarks. Our code is available at https://github.com/zyxia1009/CVPR2024-TSPNet.

📄 PDF Abstract BibTeX

Code (1)

zyxia1009/cvpr2024-tspnet 공식 구현 pytorch

Tasks

Action LocalizationTemporal Action LocalizationWeakly-supervised Temporal Action Localization

Methods 이 논문이 사용한 방법론

BASE 설명 없음
CSL Circular Smooth Label (CSL) is a classification-based rotation detection technique for arbitrary-oriented object detection. It is used for circularly distributed angle…

Similar Papers 제목 키워드 기반

SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding

2024-04-01 · Wenrui Li, Xiaopeng Hong, Ruiqin Xiong, Xiaopeng Fan

Temporal video grounding (TVG) is a critical task in video content understanding, requiring precise alignment between video content and natural language instructions. Despite significant advancements, existing methods fa…

MambaState Space ModelsVideo Grounding

SGCCNet: Single-Stage 3D Object Detector With Saliency-Guided Data Augmentation and Confidence Correction Mechanism

2024-07-01 · Ao Liang, Wenyu Chen, Jian Fang, Huaici Zhao

The single-stage point-based 3D object detectors have attracted widespread research interest due to their advantages of lightweight and fast inference speed. However, they still face challenges such as inadequate learnin…

Data Augmentation

Unsupervised Video Analysis Based on a Spatiotemporal Saliency Detector

2015-03-24 · Qiang Zhang, Yilin Wang, Baoxin Li

Visual saliency, which predicts regions in the field of view that draw the most visual attention, has attracted a lot of interest from researchers. It has already been used in several vision tasks, e.g., image classifica…

Anomaly DetectionForeground Segmentationimage-classificationImage Classification+4

Evaluating the Faithfulness of Saliency-based Explanations for Deep Learning Models for Temporal Colour Constancy

2022-11-15 · Matteo Rizzo, Cristina Conati, Daesik Jang, Hui Hu

The opacity of deep learning models constrains their debugging and improvement. Augmenting deep models with saliency-based strategies, such as attention, has been claimed to help get a better understanding of the decisio…

Decision Making

Model-guided Multi-path Knowledge Aggregation for Aerial Saliency Prediction

2018-11-14 · Kui Fu, Jia Li, Yu Zhang, Hongze Shen 외

As an emerging vision platform, a drone can look from many abnormal viewpoints which brings many new challenges into the classic vision task of video saliency prediction. To investigate these challenges, this paper propo…

Aerial Video Saliency PredictionPredictionSaliency PredictionTransfer Learning+1