A Sentence Is Worth a Thousand Pixels
We are interested in holistic scene understanding where images are accompanied with text in the form of complex sentential descriptions. We propose a holistic conditional random field model for semantic parsing which reasons jointly about which objects are present in the scene, their spatial extent as well as semantic segmentation, and employs text as well as image information as input. We automatically parse the sentences and extract objects and their relationships, and incorporate them into the model, both via potentials as well as by re-ranking candidate detections. We demonstrate the effectiveness of our approach in the challenging UIUC sentences dataset and show segmentation improvements of 12.5% over the visual only model and detection improvements of 5% AP over deformable part-based models [8].
Code (0)
등록된 구현이 없습니다.
Tasks
Re-RankingScene UnderstandingSegmentationSemantic ParsingSemantic SegmentationSentenceSimilar Papers 제목 키워드 기반
Discoverability in Satellite Imagery: A Good Sentence is Worth a Thousand Pictures
Small satellite constellations provide daily global coverage of the earth's landmass, but image enrichment relies on automating key tasks like change detection or feature searches. For example, to extract text annotation…
Change DetectionDescriptiveImage CaptioningSentence+2A Picture is Worth a Thousand Words: Language Models Plan from Pixels
Planning is an important capability of artificial agents that perform long-horizon tasks in real-world environments. In this work, we explore the use of pre-trained language models (PLMs) to reason about plan sequences f…
Neural Check-Worthiness Ranking with Weak Supervision: Finding Sentences for Fact-Checking
Automatic fact-checking systems detect misinformation, such as fake news, by (i) selecting check-worthy sentences for fact-checking, (ii) gathering related information to the sentences, and (iii) inferring the factuality…
Fact CheckingMisinformationSentenceLearning the 2-D Topology of Images
We study the following question: is the two-dimensional structure of images a very strong prior or is it something that can be learned with a few examples of natural images? If someone gave us a learning task involving i…
phi-LSTM: A Phrase-based Hierarchical LSTM Model for Image Captioning
A picture is worth a thousand words. Not until recently, however, we noticed some success stories in understanding of visual scenes: a model that is able to detect/name objects, describe their attributes, and recognize t…
Image CaptioningImage DescriptionSentence