paper-with-me

Papers

MITFAS: Mutual Information based Temporal Feature Alignment and Sampling for Aerial Video Action Recognition

2023-03-05 · Ruiqi Xian, Xijun Wang, Dinesh Manocha

We present a novel approach for action recognition in UAV videos. Our formulation is designed to handle occlusion and viewpoint changes caused by the movement of a UAV. We use the concept of mutual information to compute and align the regions corresponding to human action or motion in the temporal domain. This enables our recognition model to learn from the key features associated with the motion. We also propose a novel frame sampling method that uses joint mutual information to acquire the most informative frame sequence in UAV videos. We have integrated our approach with X3D and evaluated the performance on multiple datasets. In practice, we achieve 18.9% improvement in Top-1 accuracy over current state-of-the-art methods on UAV-Human(Li et al., 2021), 7.3% improvement on Drone-Action(Perera et al., 2019), and 7.16% improvement on NEC Drones(Choi et al., 2020).

📄 PDF Abstract BibTeX arXiv:2303.02575

Code (1)

ricky-xian/mitfas 공식 구현

Tasks

Action RecognitionTemporal Action Localization

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Zero-shot Skeleton-based Action Recognition via Mutual Information Estimation and Maximization

2023-08-07 · Yujie Zhou, Wenwen Qiang, Anyi Rao, Ning Lin 외

Zero-shot skeleton-based action recognition aims to recognize actions of unseen categories after training on data of seen categories. The key is to build the connection between visual and semantic space from seen to unse…

Action RecognitionMutual Information EstimationSkeleton Based Action RecognitionZero Shot Skeletal Action Recognition+1

Mutually-paced Knowledge Distillation for Cross-lingual Temporal Knowledge Graph Reasoning

2023-03-27 · Ruijie Wang, Zheng Li, Jingfeng Yang, Tianyu Cao 외

This paper investigates cross-lingual temporal knowledge graph reasoning problem, which aims to facilitate reasoning on Temporal Knowledge Graphs (TKGs) in low-resource languages by transfering knowledge from TKGs in hig…

Knowledge DistillationKnowledge GraphsTransfer Learning

STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing

2025-03-29 · Zijun Ding, Mingdie Xiong, Congcong Zhu, Jingrun Chen

Existing audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly degrades the synthesis stability for the d…

Implicit Temporal Modeling with Learnable Alignment for Video Recognition

2023-04-20 · ICCV 2023 1 · Shuyuan Tu, Qi Dai, Zuxuan Wu, Zhi-Qi Cheng 외

Contrastive language-image pretraining (CLIP) has demonstrated remarkable success in various image tasks. However, how to extend CLIP with effective temporal modeling is still an open and crucial problem. Existing factor…

Action ClassificationAction RecognitionVideo Recognition

Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose Estimation

2022-03-29 · CVPR 2022 1 · Zhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu 외

Multi-frame human pose estimation has long been a compelling and fundamental problem in computer vision. This task is challenging due to fast motion and pose occlusion that frequently occur in videos. State-of-the-art me…

Pose Estimation