A Multi-Modal Transformer Network for Action Detection
This paper proposes a novel multi-modal transformer network for detecting actions in untrimmed videos. To enrich the action features, our transformer network utilizes a new multi-modal attention mechanism that computes the correlations between different spatial and motion modalities combinations. Exploring such correlations for actions has not been attempted previously. To use the motion and spatial modality more effectively, we suggest an algorithm that corrects the motion distortion caused by camera movement. Such motion distortion, common in untrimmed videos, severely reduces the expressive power of motion features such as optical flow fields. Our proposed algorithm outperforms the state-of-the-art methods on two public benchmarks, THUMOS14 and ActivityNet. We also conducted comparative experiments on our new instructional activity dataset, including a large set of challenging classroom videos captured from elementary schools.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionOptical Flow EstimationSimilar Papers 제목 키워드 기반
Cross-Modality Fusion Transformer for Multispectral Object Detection
Multispectral image pairs can provide the combined information, making object detection applications more reliable and robust in the open world. To fully exploit the different modalities, we present a simple yet effectiv…
Multispectral Object DetectionObjectobject-detectionObject Detection+1Multi-Modal Learning for AU Detection Based on Multi-Head Fused Transformers
Multi-modal learning has been intensified in recent years, especially for applications in facial analysis and action unit detection whilst there still exist two main challenges in terms of 1) relevant feature learning fo…
Action Unit DetectionMDD-Net: Multimodal Depression Detection through Mutual Transformer
Depression is a major mental health condition that severely impacts the emotional and physical well-being of individuals. The simple nature of data collection from social media platforms has attracted significant interes…
SeaDATE: Remedy Dual-Attention Transformer with Semantic Alignment via Contrast Learning for Multimodal Object Detection
Multimodal object detection leverages diverse modal information to enhance the accuracy and robustness of detectors. By learning long-term dependencies, Transformer can effectively integrate multimodal features in the fe…
Contrastive Learningobject-detectionObject DetectionDual Sparse Aggregation Transformer for Multispectral Object Detection
Transformer-based approaches have obtained excellent performance in multispectral object detection tasks due to their ability to model long-range dependencies and capture complementary information. However, previous tran…
Multispectral Object Detection