AssembleNet++: Assembling Modality Representations via Attention Connections - Supplementary Material -
We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the importance of features at each convolutional block of the network. A new network component named peer-attention is introduced, which dynamically learns the attention weights using another block or input modality. Even without pre-training, our models outperform the previous work on standard public activity recognition datasets with continuous videos, establishing new state-of-the-art. We also confirm that our findings of having neural connections from the object modality and the use of peer-attention is generally applicable for different existing architectures, improving their performances. We name our model explicitly as AssembleNet++. The code will be available at: https://sites.google.com/corp/view/assemblenet/
Code (0)
등록된 구현이 없습니다.
Tasks
Activity RecognitionSimilar Papers 제목 키워드 기반
AssembleNet++: Assembling Modality Representations via Attention Connections
We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the imp…
Action ClassificationActivity RecognitionReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco Reconstruction
The task of reassembly is a significant challenge across multiple domains, including archaeology, genomics, and molecular docking, requiring the precise placement and orientation of elements to reconstruct an original st…
Molecular DockingPose EstimationAssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures
Learning to represent videos is a very challenging task both algorithmically and computationally. Standard video CNN architectures have been designed by directly extending architectures devised for image understanding to…
Action ClassificationAction RecognitionMultimodal Activity RecognitionOptical Flow Estimation+2Disassembling Object Representations without Labels
In this paper, we study a new representation-learning task, which we termed as disassembling object representations. Given an image featuring multiple objects, the goal of disassembling is to acquire a latent representat…
General ClassificationGenerative Adversarial NetworkObjectRepresentation Learning+1NestedFormer: Nested Modality-Aware Transformer for Brain Tumor Segmentation
Multi-modal MR imaging is routinely used in clinical practice to diagnose and investigate brain tumors by providing rich complementary information. Previous multi-modal MRI segmentation methods usually perform modal fusi…
Brain Tumor SegmentationDecoderMRI segmentationSegmentation+1