paper-with-me

Papers

Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks

2022-11-03 · Matthew Kowal, Mennatullah Siam, Md Amirul Islam, Neil D. B. Bruce, Richard P. Wildes, Konstantinos G. Derpanis

There is limited understanding of the information captured by deep spatiotemporal models in their intermediate representations. For example, while evidence suggests that action recognition algorithms are heavily influenced by visual appearance in single frames, no quantitative methodology exists for evaluating such static bias in the latent representation compared to bias toward dynamics. We tackle this challenge by proposing an approach for quantifying the static and dynamic biases of any spatiotemporal model, and apply our approach to three tasks, action recognition, automatic video object segmentation (AVOS) and video instance segmentation (VIS). Our key findings are: (i) Most examined models are biased toward static information. (ii) Some datasets that are assumed to be biased toward dynamics are actually biased toward static information. (iii) Individual channels in an architecture can be biased toward static, dynamic or a combination of the two. (iv) Most models converge to their culminating biases in the first half of training. We then explore how these biases affect performance on dynamically biased datasets. For action recognition, we propose StaticDropout, a semantically guided dropout that debiases a model from static information toward dynamics. For AVOS, we design a better combination of fusion and cross connection layers compared with previous architectures.

📄 PDF Abstract BibTeX arXiv:2211.01783

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionInstance SegmentationSemantic SegmentationVideo Instance SegmentationVideo Object SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic Information

2022-06-06 · CVPR 2022 1 · Matthew Kowal, Mennatullah Siam, Md Amirul Islam, Neil D. B. Bruce 외

Deep spatiotemporal models are used in a variety of computer vision tasks, such as action recognition and video object segmentation. Currently, there is a limited understanding of what information is captured by these mo…

Action RecognitionSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Cell Behavior Video Classification Challenge, a benchmark for computer vision methods in time-lapse microscopy

2026-01-15 · Raffaella Fiamma Cabini, Deborah Barkauskas, Guangyu Chen, Zhi-Qi Cheng 외 arxiv

The classification of microscopy videos capturing complex cellular behaviors is crucial for understanding and quantifying the dynamics of biological processes over time. However, it remains a frontier in computer vision,…

Video Classification

Link the World: Improving Open-domain Conversation with Dynamic Spatiotemporal-aware Knowledge

2022-06-28 · Han Zhou, Xinchao Xu, Wenquan Wu, Zheng-Yu Niu 외

Making chatbots world aware in a conversation like a human is a crucial challenge, where the world may contain dynamic knowledge and spatiotemporal state. Several recent advances have tried to link the dialog system to a…

Informativeness

Energy Prediction using Spatiotemporal Pattern Networks

2017-02-03 · Zhanhong Jiang, Chao Liu, Adedotun Akintayo, Gregor Henze 외

This paper presents a novel data-driven technique based on the spatiotemporal pattern network (STPN) for energy/power prediction for complex dynamical systems. Built on symbolic dynamic filtering, the STPN framework is u…

Prediction

STVGFormer: Spatio-Temporal Video Grounding with Static-Dynamic Cross-Modal Understanding

2022-07-06 · Zihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin 외

In this technical report, we introduce our solution to human-centric spatio-temporal video grounding task. We propose a concise and effective framework named STVGFormer, which models spatiotemporal visual-linguistic depe…

Spatio-Temporal Video GroundingVideo Grounding