Segment-Phrase Table for Semantic Segmentation, Visual Entailment and Paraphrasing
We introduce Segment-Phrase Table (SPT), a large collection of bijective associations between textual phrases and their corresponding segmentations. Leveraging recent progress in object recognition and natural language semantics, we show how we can successfully build a high-quality segment-phrase table using minimal human supervision. More importantly, we demonstrate the unique value unleashed by this rich bimodal resource, for both vision as well as natural language understanding. First, we show that fine-grained textual labels facilitate contextual reasoning that helps in satisfying semantic constraints across image segments. This feature enables us to achieve state-of-the-art segmentation results on benchmark datasets. Next, we show that the association of high-quality segmentations to textual phrases aids in richer semantic understanding and reasoning of these textual phrases. Leveraging this feature, we motivate the problem of visual entailment and visual paraphrasing, and demonstrate its utility on a large dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language UnderstandingObject RecognitionSemantic SegmentationVisual EntailmentSimilar Papers 제목 키워드 기반
PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative Grounding
Panoptic Narrative Grounding (PNG) is an emerging task whose goal is to segment visual objects of things and stuff categories described by dense narrative captions of a still image. The previous two-stage approach first …
Panoptic SegmentationSegmentationSemantic correspondenceSegmentation from Natural Language Expressions
In this paper we approach the novel problem of segmenting an image based on a natural language expression. This is different from traditional semantic segmentation over a predefined set of semantic classes, as e.g., the …
Referring Expression SegmentationSegmentationSemantic SegmentationPanoptic Narrative Grounding
This paper proposes Panoptic Narrative Grounding, a spatially fine and general formulation of the natural language visual grounding problem. We establish an experimental framework for the study of this new task, includin…
Natural Language Visual GroundingPanoptic SegmentationVisual GroundingPhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
Understanding how natural language phrases correspond to specific regions in images is a key challenge in multimodal semantic segmentation. Recent advances in phrase grounding are largely limited to single-view images, n…
Semantic SegmentationImage SegmentationPhrase GroundingPanoptic Narrative Grounding
This paper proposes Panoptic Narrative Grounding, a spatially fine and general formulation of the natural language visual grounding problem. We establish an experimental framework for the study of this new task, incl…
Natural Language Visual GroundingPanoptic SegmentationVisual Grounding