paper-with-me

홈 › Papers

Exploring Explainability in Video Action Recognition

2024-04-13 · Avinab Saha, Shashank Gupta, Sravan Kumar Ankireddy, Karl Chahine, Joydeep Ghosh

Image Classification and Video Action Recognition are perhaps the two most foundational tasks in computer vision. Consequently, explaining the inner workings of trained deep neural networks is of prime importance. While numerous efforts focus on explaining the decisions of trained deep neural networks in image classification, exploration in the domain of its temporal version, video action recognition, has been scant. In this work, we take a deeper look at this problem. We begin by revisiting Grad-CAM, one of the popular feature attribution methods for Image Classification, and its extension to Video Action Recognition tasks and examine the method's limitations. To address these, we introduce Video-TCAV, by building on TCAV for Image Classification tasks, which aims to quantify the importance of specific concepts in the decision-making process of Video Action Recognition models. As the scalable generation of concepts is still an open problem, we propose a machine-assisted approach to generate spatial and spatiotemporal concepts relevant to Video Action Recognition for testing Video-TCAV. We then establish the importance of temporally-varying concepts by demonstrating the superiority of dynamic spatiotemporal concepts over trivial spatial concepts. In conclusion, we introduce a framework for investigating hypotheses in action recognition and quantitatively testing them, thus advancing research in the explainability of deep neural networks used in video action recognition.

📄 PDF Abstract BibTeX arXiv:2404.09067

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionClassificationimage-classificationImage ClassificationTemporal Action Localization

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Multi-velocity neural networks for gesture recognition in videos

2016-03-22 · Otkrist Gupta, Dan Raviv, Ramesh Raskar

We present a new action recognition deep neural network which adaptively learns the best action velocities in addition to the classification. While deep neural networks have reached maturity for image understanding tasks…

Action RecognitionGeneral ClassificationGesture RecognitionTemporal Action Localization

Exploring Temporal Information for Improved Video Understanding

2019-05-25 · Yi Zhu

In this dissertation, I present my work towards exploring temporal information for better video understanding. Specifically, I have worked on two problems: action recognition and semantic segmentation. For action recogni…

Action RecognitionOptical Flow EstimationSegmentationSemantic Segmentation+3

Depth2Action: Exploring Embedded Depth for Large-Scale Action Recognition

2016-08-15 · Yi Zhu, Shawn Newsam

This paper performs the first investigation into depth for large-scale human action recognition in video where the depth cues are estimated from the videos themselves. We develop a new framework called depth2action and e…

Action RecognitionTemporal Action Localization

Mutual Context Network for Jointly Estimating Egocentric Gaze and Actions

2019-01-07 · Yifei Huang, Zhenqiang Li, Minjie Cai, Yoichi Sato

In this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context. Our assumption is that in the procedure of performing a manipulation task, what…

Action RecognitionGaze PredictionPredictionTemporal Action Localization

Generating Action-conditioned Prompts for Open-vocabulary Video Action Recognition

2023-12-04 · Chengyou Jia, Minnan Luo, Xiaojun Chang, Zhuohang Dang 외

Exploring open-vocabulary video action recognition is a promising venture, which aims to recognize previously unseen actions within any arbitrary set of categories. Existing methods typically adapt pretrained image-text …

Action RecognitionDescriptiveTemporal Action Localization