paper-with-me

홈 › Papers

Harvesting Information from Captions for Weakly Supervised Semantic Segmentation

2019-05-16 · Johann Sawatzky, Debayan Banerjee, Juergen Gall

Since acquiring pixel-wise annotations for training convolutional neural networks for semantic image segmentation is time-consuming, weakly supervised approaches that only require class tags have been proposed. In this work, we propose another form of supervision, namely image captions as they can be found on the Internet. These captions have two advantages. They do not require additional curation as it is the case for the clean class tags used by current weakly supervised approaches and they provide textual context for the classes present in an image. To leverage such textual context, we deploy a multi-modal network that learns a joint embedding of the visual representation of the image and the textual representation of the caption. The network estimates text activation maps (TAMs) for class names as well as compound concepts, i.e. combinations of nouns and their attributes. The TAMs of compound concepts describing classes of interest substantially improve the quality of the estimated class activation maps which are then used to train a network for semantic segmentation. We evaluate our method on the COCO dataset where it achieves state of the art results for weakly supervised image segmentation.

📄 PDF Abstract BibTeX arXiv:1905.06784

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningImage SegmentationSegmentationSemantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic Segmentation

Similar Papers 제목 키워드 기반

Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection

2021-12-01 · NeurIPS 2021 12 · Jingjing Li, Wei Ji, Qi Bi, Cheng Yan 외

Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection…

object-detectionObject DetectionRGB-D Salient Object DetectionSaliency Detection+1

SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

2026-03-05 · Ye-Chan Kim, SeungJu Cha, Si-Woo Kim, Minju Jeon 외 arxiv

Weakly-Supervised Dense Video Captioning aims to localize and describe events in videos trained only on caption annotations, without temporal boundaries. Prior work introduced an implicit supervision paradigm based on Ga…

Dense Video Captioning

Contrastive Learning for Weakly Supervised Phrase Grounding

2020-06-17 · ECCV 2020 8 · Tanmay Gupta, Arash Vahdat, Gal Chechik, Xiaodong Yang 외

Phrase grounding, the problem of associating image regions to caption words, is a crucial component of vision-language tasks. We show that phrase grounding can be learned by optimizing word-region attention to maximize a…

Contrastive LearningLanguage ModelingLanguage ModellingPhrase Grounding

LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation

2023-10-16 · CVPR 2024 1 · Kibum Kim, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In 외

Weakly-Supervised Scene Graph Generation (WSSGG) research has recently emerged as an alternative to the fully-supervised approach that heavily relies on costly annotations. In this regard, studies on WSSGG have utilized …

Few-Shot LearningLarge Language ModelScene Graph GenerationTriplet+1

Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding

2022-10-21 · Findings (EMNLP) 2021 11 · Yuechen Wang, Wengang Zhou, Houqiang Li

Temporal language grounding (TLG) aims to localize a video segment in an untrimmed video based on a natural language description. To alleviate the expensive cost of manual annotations for temporal boundary labels, we are…

cross-modal alignmentSentence