Visual-Semantic Scene Understanding by Sharing Labels in a Context Network
We consider the problem of naming objects in complex, natural scenes containing widely varying object appearance and subtly different names. Informed by cognitive research, we propose an approach based on sharing context based object hypotheses between visual and lexical spaces. To this end, we present the Visual Semantic Integration Model (VSIM) that represents object labels as entities shared between semantic and visual contexts and infers a new image by updating labels through context switching. At the core of VSIM is a semantic Pachinko Allocation Model and a visual nearest neighbor Latent Dirichlet Allocation Model. For inference, we derive an iterative Data Augmentation algorithm that pools the label probabilities and maximizes the joint label posterior of an image. Our model surpasses the performance of state-of-art methods in several visual tasks on the challenging SUN09 dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationObjectScene UnderstandingSimilar Papers 제목 키워드 기반
Semantically-aware Neural Radiance Fields for Visual Scene Understanding: A Comprehensive Review
This review thoroughly examines the role of semantically-aware Neural Radiance Fields (NeRFs) in visual scene understanding, covering an analysis of over 250 scholarly papers. It explores how NeRFs adeptly infer 3D repre…
Panoptic SegmentationScene SegmentationScene UnderstandingSegmentationIncorporating Scene Context and Semantic Labels for Enhanced Group-level Emotion Recognition
Group-level emotion recognition (GER) aims to identify holistic emotions within a scene involving multiple individuals. Current existed methods underestimate the importance of visual scene contextual information in model…
Emotion RecognitionIn-Place Scene Labelling and Understanding with Implicit Scene Representation
Semantic labelling is highly correlated with geometry and radiance reconstruction, as scene entities with similar shape and appearance are more likely to come from similar classes. Recent implicit neural reconstruction t…
DenoisingNeRFSuper-ResolutionStableSemantics: A Synthetic Language-Vision Dataset of Semantic Representations in Naturalistic Images
Understanding the semantics of visual scenes is a fundamental challenge in Computer Vision. A key aspect of this challenge is that objects sharing similar semantic meanings or functions can exhibit striking visual differ…
Object RecognitionScene UnderstandingCityscapes-Panoptic-Parts and PASCAL-Panoptic-Parts datasets for Scene Understanding
In this technical report, we present two novel datasets for image scene understanding. Both datasets have annotations compatible with panoptic segmentation and additionally they have part-level labels for selected semant…
Human Part SegmentationPanoptic SegmentationPart-aware Panoptic SegmentationScene Understanding+2