AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition
Multi-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learning offers excellent recognition results, its computational expense limits its impact for many real-world applications. In this paper, we propose an adaptive multi-modal learning framework, called AdaMML, that selects on-the-fly the optimal modalities for each segment conditioned on the input for efficient video recognition. Specifically, given a video segment, a multi-modal policy network is used to decide what modalities should be used for processing by the recognition model, with the goal of improving both accuracy and efficiency. We efficiently train the policy network jointly with the recognition model using standard back-propagation. Extensive experiments on four challenging diverse datasets demonstrate that our proposed adaptive approach yields 35%-55% reduction in computation when compared to the traditional baseline that simply uses all the modalities irrespective of the input, while also achieving consistent improvements in accuracy over the state-of-the-art methods.
Code (1)
Tasks
Video RecognitionSimilar Papers 제목 키워드 기반
Frame Aggregation and Multi-Modal Fusion Framework for Video-Based Person Recognition
Video-based person recognition is challenging due to persons being blocked and blurred, and the variation of shooting angle. Previous research always focused on person recognition on still images, ignoring similarity and…
Person RecognitionHCMS: Hierarchical and Conditional Modality Selection for Efficient Video Recognition
Videos are multimodal in nature. Conventional video recognition pipelines typically fuse multimodal features for improved performance. However, this is not only computationally expensive but also neglects the fact that d…
Video RecognitionLearning Cross-modal Contrastive Features for Video Domain Adaptation
Learning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature …
Action RecognitionContrastive LearningDomain AdaptationOptical Flow EstimationInteract Before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action Recognition
Unsupervised domain adaptive video action recognition aims to recognize actions of a target domain using a model trained with only out-of-domain (source) annotations. The inherent complexity of videos makes this task…
Action RecognitionTemporal Action LocalizationHEU Emotion: A Large-scale Database for Multi-modal Emotion Recognition in the Wild
The study of affective computing in the wild setting is underpinned by databases. Existing multimodal emotion databases in the real-world conditions are few and small, with a limited number of subjects and expressed in a…
Emotion RecognitionFacial Expression RecognitionFacial Expression Recognition (FER)