paper-with-me

홈 › Papers

Delving Deep Into Many-to-Many Attention for Few-Shot Video Object Segmentation

2021-06-19 · CVPR 2021 1 · Haoxin Chen, Hanjie Wu, Nanxuan Zhao, Sucheng Ren, Shengfeng He

This paper tackles the task of Few-Shot Video Object Segmentation (FSVOS), i.e., segmenting objects in the query videos with certain class specified in a few labeled support images. The key is to model the relationship between the query videos and the support images for propagating the object information. This is a many-to-many problem and often relies on full-rank attention, which is computationally intensive. In this paper, we propose a novel Domain Agent Network (DAN), breaking down the full-rank attention into two smaller ones. We consider one single frame of the query video as the domain agent, bridging between the support images and the query video. Our DAN allows a linear space and time complexity as opposed to the original quadratic form with no loss of performance. In addition, we introduce a learning strategy by combining meta-learning with online learning to further improve the segmentation accuracy. We build a FSVOS benchmark on the Youtube-VIS dataset and conduct experiments to demonstrate that our method outperforms baselines on both computational cost and accuracy, achieving the state-of-the-art performance.

📄 PDF Abstract BibTeX

Code (1)

scutpaul/DANet 공식 구현 pytorch

Tasks

Meta-LearningSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

Hubness and Pollution: Delving into Cross-Space Mapping for Zero-Shot Learning

2015-07-01 · IJCNLP 2015 7 · Angeliki Lazaridou, Georgiana Dinu, Marco Baroni
Semantic Textual SimilarityZero-Shot Learning

Delving into Differentially Private Transformer

2024-05-28 · Youlong Ding, Xueyang Wu, Yining Meng, Yonggang Luo 외

Deep learning with differential privacy (DP) has garnered significant attention over the past years, leading to the development of numerous methods aimed at enhancing model accuracy and training efficiency. This paper de…

Enhancing Zero-Shot Many to Many Voice Conversion with Self-Attention VAE

2022-03-30 · Ziang Long, Yunling Zheng, Meng Yu, Jack Xin

Variational auto-encoder (VAE) is an effective neural network architecture to disentangle a speech utterance into speaker identity and linguistic content latent embeddings, then generate an utterance for a target speaker…

DecoderSentenceVoice Conversion

Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention

2025-03-11 · Emily Xiao, Chin-Jou Li, Yilin Zhang, Graham Neubig 외

Many-shot in-context learning has recently shown promise as an alternative to finetuning, with the major advantage that the same model can be served for multiple tasks. However, this shifts the computational burden from …

In-Context LearningRetrieval

Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning

2024-06-21 · Brandon Huang, Chancharik Mitra, Assaf Arbelle, Leonid Karlinsky 외

The recent success of interleaved Large Multimodal Models (LMMs) in few-shot learning suggests that in-context learning (ICL) with many examples can be promising for learning new tasks. However, this many-shot multimodal…

Few-Shot LearningIn-Context Learning