paper-with-me

Papers

AssembleNet++: Assembling Modality Representations via Attention Connections - Supplementary Material -

2020-08-01 · ECCV 2020 8 · Michael S. Ryoo, AJ Piergiovanni, Juhana Kangaspunta, Anelia Angelova

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the importance of features at each convolutional block of the network. A new network component named peer-attention is introduced, which dynamically learns the attention weights using another block or input modality. Even without pre-training, our models outperform the previous work on standard public activity recognition datasets with continuous videos, establishing new state-of-the-art. We also confirm that our findings of having neural connections from the object modality and the use of peer-attention is generally applicable for different existing architectures, improving their performances. We name our model explicitly as AssembleNet++. The code will be available at: https://sites.google.com/corp/view/assemblenet/

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Activity Recognition

Similar Papers 제목 키워드 기반

AssembleNet++: Assembling Modality Representations via Attention Connections

2020-08-18 · Michael S. Ryoo, AJ Piergiovanni, Juhana Kangaspunta, Anelia Angelova

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the imp…

Action ClassificationActivity Recognition

ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco Reconstruction

2025-05-27 · Adeela Islam, Stefano Fiorini, Stuart James, Pietro Morerio 외

The task of reassembly is a significant challenge across multiple domains, including archaeology, genomics, and molecular docking, requiring the precise placement and orientation of elements to reconstruct an original st…

Molecular DockingPose Estimation

AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures

2019-05-30 · ICLR 2020 1 · Michael S. Ryoo, AJ Piergiovanni, Mingxing Tan, Anelia Angelova

Learning to represent videos is a very challenging task both algorithmically and computationally. Standard video CNN architectures have been designed by directly extending architectures devised for image understanding to…

Action ClassificationAction RecognitionMultimodal Activity RecognitionOptical Flow Estimation+2

Disassembling Object Representations without Labels

2020-04-03 · Zunlei Feng, Xinchao Wang, Yongming He, Yike Yuan 외

In this paper, we study a new representation-learning task, which we termed as disassembling object representations. Given an image featuring multiple objects, the goal of disassembling is to acquire a latent representat…

General ClassificationGenerative Adversarial NetworkObjectRepresentation Learning+1

NestedFormer: Nested Modality-Aware Transformer for Brain Tumor Segmentation

2022-08-31 · Zhaohu Xing, Lequan Yu, Liang Wan, Tong Han 외

Multi-modal MR imaging is routinely used in clinical practice to diagnose and investigate brain tumors by providing rich complementary information. Previous multi-modal MRI segmentation methods usually perform modal fusi…

Brain Tumor SegmentationDecoderMRI segmentationSegmentation+1