DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension
In this paper, we focus on weakly supervised referring expression comprehension (REC), and identify that the lack of fine-grained visual capability greatly limits the upper performance bound of existing methods. To address this issue, we propose a novel framework for weakly supervised REC, namely Dynamic Visual routing Network (DViN), which overcomes the visual shortcomings from the perspective of feature combination and alignment. In particular, DViN is equipped with a novel sparse routing mechanism to efficiently combine features of multiple visual encoders in a dynamic manner, thus improving the visual descriptive power. Besides, we further propose an innovative weakly supervised objective, namely Routing-based Feature Alignment (RFA), which facilitates the visual understanding of routed features through the intra-modal and inter-modal alignment. To validate DViN, we conduct extensive experiments on four REC benchmark datasets. Experiments demonstrate that DViN achieves state-of-the-art results on four benchmarks while maintaining competitive inference efficiency. Besides, the strong generalization ability of DViN is also validated on weakly supervised referring expression segmentation. Source codes are anonymously released at: https://anonymous.4open.science/r/DViN-7736.
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveReferring ExpressionReferring Expression ComprehensionReferring Expression SegmentationWeakly Supervised Referring Expression SegmentationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Empowering Collaborative Filtering with Principled Adversarial Contrastive Loss
Contrastive Learning (CL) has achieved impressive performance in self-supervised learning tasks, showing superior generalization ability. Inspired by the success, adopting CL into collaborative filtering (CF) is prevaili…
Collaborative FilteringContrastive LearningData AugmentationRecommendation Systems+1Imperceptible Adversarial Attack via Invertible Neural Networks
Adding perturbations via utilizing auxiliary gradient information or discarding existing details of the benign images are two common approaches for generating adversarial examples. Though visual imperceptibility is the d…
Adversarial AttackAutomatic Discovery of Novel Intents & Domains from Text Utterances
One of the primary tasks in Natural Language Understanding (NLU) is to recognize the intents as well as domains of users' spoken and written language utterances. Most existing research formulates this as a supervised cla…
General ClassificationNatural Language UnderstandingTransfer LearningWeather-Conditioned Branch Routing for Robust LiDAR-Radar 3D Object Detection
Robust 3D object detection in adverse weather is highly challenging due to the varying reliability of different sensors. While existing LiDAR-4D radar fusion methods improve robustness, they predominantly rely on fixed o…
Robust 3D Object DetectionWeakly Supervised Visual Semantic Parsing
Scene Graph Generation (SGG) aims to extract entities, predicates and their semantic structure from images, enabling deep understanding of visual content, with many applications such as visual reasoning and image retriev…
Graph GenerationImage RetrievalRetrievalScene Graph Generation+3