Scene Recognition with Prototype-agnostic Scene Layout
Abstract--- Exploiting the spatial structure in scene images is a key research direction for scene recognition. Due to the large intra-class structural diversity, building and modeling flexible structural layout to adapt various image characteristics is a challenge. Existing structural modeling methods in scene recognition either focus on predefined grids or rely on learned prototypes, which all have limited representative ability. In this paper, we propose Prototype-agnostic Scene Layout (PaSL) construction method to build the spatial structure for each image without conforming to any prototype. Our PaSL can flexibly capture the diverse spatial characteristic of scene images and have considerable generalization capability. Given a PaSL, we build Layout Graph Network (LGN) where regions in PaSL are defined as nodes and two kinds of independent relations between regions are encoded as edges. The LGN aims to incorporate two topological structures (formed in spatial and semantic similarity dimensions) into image representations through graph convolution. Extensive experiments show that our approach achieves state-of-the-art results on widely recognized MIT67 and SUN397 datasets without multi-model or multi-scale fusion. Moreover, we also conduct the experiments on one of the largest scale datasets, Places365. The results demonstrate the proposed method can be well generalized and obtains competitive performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Scene RecognitionSemantic SimilaritySemantic Textual SimilaritySimilar Papers 제목 키워드 기반
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-bas…
DiversityImage GenerationInstance SegmentationLayout Generation+2Aerial Scene Understanding in The Wild: Multi-Scene Recognition via Prototype-based Memory Networks
Aerial scene recognition is a fundamental visual task and has attracted an increasing research interest in the last few years. Most of current researches mainly deploy efforts to categorize an aerial image into one scene…
RetrievalScene RecognitionScene UnderstandingInter-object Discriminative Graph Modeling for Indoor Scene Recognition
Variable scene layouts and coexisting objects across scenes make indoor scene recognition still a challenging task. Leveraging object information within scenes to enhance the distinguishability of feature representations…
ObjectScene RecognitionSemantic-embedded Similarity Prototype for Scene Recognition
Due to the high inter-class similarity caused by the complex composition and the co-existing objects across scenes, numerous studies have explored object semantic knowledge within scenes to improve scene recognition. How…
Objectobject-detectionObject DetectionScene Recognition+1IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning
Spatial question answering is the dominant paradigm for evaluating spatial intelligence in Vision-Language Models (VLMs), but it leaves a complementary axis of spatial competence under-evaluated: holistic 3D layout infer…
Object RecognitionQuestion AnsweringSpatial Reasoning