paper-with-me

홈 › Papers

End-to-end weakly-supervised semantic alignment

2017-12-19 · CVPR 2018 6 · Ignacio Rocco, Relja Arandjelović, Josef Sivic

We tackle the task of semantic alignment where the goal is to compute dense semantic correspondence aligning two images depicting objects of the same category. This is a challenging task due to large intra-class variation, changes in viewpoint and background clutter. We present the following three principal contributions. First, we develop a convolutional neural network architecture for semantic alignment that is trainable in an end-to-end manner from weak image-level supervision in the form of matching image pairs. The outcome is that parameters are learnt from rich appearance variation present in different but semantically related images without the need for tedious manual annotation of correspondences at training time. Second, the main component of this architecture is a differentiable soft inlier scoring module, inspired by the RANSAC inlier scoring procedure, that computes the quality of the alignment based on only geometrically consistent correspondences thereby reducing the effect of background clutter. Third, we demonstrate that the proposed approach achieves state-of-the-art performance on multiple standard benchmarks for semantic alignment.

📄 PDF Abstract BibTeX arXiv:1712.06861

Code (2)

hukim1112/weakalign pytorch
ignacio-rocco/weakalign pytorch

Tasks

Semantic correspondence

Similar Papers 제목 키워드 기반

Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding

2022-10-21 · Findings (EMNLP) 2021 11 · Yuechen Wang, Wengang Zhou, Houqiang Li

Temporal language grounding (TLG) aims to localize a video segment in an untrimmed video based on a natural language description. To alleviate the expensive cost of manual annotations for temporal boundary labels, we are…

cross-modal alignmentSentence

Distribution Guidance Network for Weakly Supervised Point Cloud Semantic Segmentation

2024-10-10 · Zhiyi Pan, Wei Gao, Shan Liu, Ge Li

Despite alleviating the dependence on dense annotations inherent to fully supervised methods, weakly supervised point cloud semantic segmentation suffers from inadequate supervision signals. In response to this challenge…

Semantic SegmentationWeakly-supervised Learning

AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding

2025-08-05 · Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao 외 arxiv

Weakly supervised visual grounding (VG) aims to locate objects in images based on text descriptions. Despite significant progress, existing methods lack strong cross-modal reasoning to distinguish subtle semantic differe…

Contrastive LearningVisual Grounding

SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised Segmentation

2025-12-01 · Xiuli Bi, Die Xiao, Junchao Fan, Bin Xiao arxiv

In recent years, Contrastive Language-Image Pretraining (CLIP) has been widely applied to Weakly Supervised Semantic Segmentation (WSSS) tasks due to its powerful cross-modal semantic understanding capabilities. This pap…

Semantic SegmentationContrastive Learning

Weakly Supervised Temporal Adjacent Network for Language Grounding

2021-06-30 · Yuechen Wang, Jiajun Deng, Wengang Zhou, Houqiang Li

Temporal language grounding (TLG) is a fundamental and challenging problem for vision and language understanding. Existing methods mainly focus on fully supervised setting with temporal boundary labels for training, whic…

Multiple Instance LearningSentence