paper-with-me

Papers

Class Feature Pyramids for Video Explanation

2019-09-18 · Alexandros Stergiou, Georgios Kapidis, Grigorios Kalliatakis, Christos Chrysoulas, Ronald Poppe, Remco Veltkamp

Deep convolutional networks are widely used in video action recognition. 3D convolutions are one prominent approach to deal with the additional time dimension. While 3D convolutions typically lead to higher accuracies, the inner workings of the trained models are more difficult to interpret. We focus on creating human-understandable visual explanations that represent the hierarchical parts of spatio-temporal networks. We introduce Class Feature Pyramids, a method that traverses the entire network structure and incrementally discovers kernels at different network depths that are informative for a specific class. Our method does not depend on the network's architecture or the type of 3D convolutions, supporting grouped and depth-wise convolutions, convolutions in fibers, and convolutions in branches. We demonstrate the method on six state-of-the-art 3D convolution neural networks (CNNs) on three action recognition (Kinetics-400, UCF-101, and HMDB-51) and two egocentric action recognition datasets (EPIC-Kitchens and EGTEA Gaze+).

📄 PDF Abstract BibTeX arXiv:1909.08611

Code (1)

alexandrosstergiou/Class_Feature_Visualization_Pyramid 공식 구현 pytorch

Tasks

Action RecognitionTemporal Action Localization

Methods 이 논문이 사용한 방법론

3D Convolution A 3D Convolution is a type of convolution where the kernel slides in 3 dimensions as opposed to 2 dimensions with 2D…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Visualizing the Passage of Time with Video Temporal Pyramids

2022-08-25 · Melissa E. Swift, Wyatt Ayers, Sophie Pallanck, Scott Wehrwein

What can we learn about a scene by watching it for months or years? A video recorded over a long timespan will depict interesting phenomena at multiple timescales, but identifying and viewing them presents a challenge. T…

3D Feature Pyramid Attention Module for Robust Visual Speech Recognition

2018-10-15 · Jing-Yun Xiao

Visual speech recognition is the task to decode the speech content from a video based on visual information, especially the movements of lips. It is also referenced as lipreading. Motivated by two problems existing in li…

LipreadingSentencespeech-recognitionSpeech Recognition+1

Temporal Action Localization with Multi-temporal Scales

2022-08-16 · Zan Gao, Xinglei Cui, Tao Zhuo, Zhiyong Cheng 외

Temporal action localization plays an important role in video analysis, which aims to localize and classify actions in untrimmed videos. The previous methods often predict actions on a feature space of a single-temporal …

Action ClassificationAction LocalizationAvgTemporal Action Localization

EMface: Detecting Hard Faces by Exploring Receptive Field Pyraminds

2021-05-21 · Leilei Cao, Yao Xiao, Lin Xu

Scale variation is one of the most challenging problems in face detection. Modern face detectors employ feature pyramids to deal with scale variation. However, it might break the feature consistency across different scal…

Face Detection

ProjectedEx: Enhancing Generation in Explainable AI for Prostate Cancer

2025-01-02 · Xuyin Qi, Zeyu Zhang, Aaron Berliano Handoko, Huazhan Zheng 외

Prostate cancer, a growing global health concern, necessitates precise diagnostic tools, with Magnetic Resonance Imaging (MRI) offering high-resolution soft tissue imaging that significantly enhances diagnostic accuracy.…

AttributeDiagnosticImage GenerationLesion Classification+1