paper-with-me

홈 › Papers

Bounding Boxes Are All We Need: Street View Image Classification via Context Encoding of Detected Buildings

2020-10-03 · Kun Zhao, Yongkun Liu, Siyuan Hao, Shaoxing Lu, Hongbin Liu, Lijian Zhou

Street view images classification aiming at urban land use analysis is difficult because the class labels (e.g., commercial area), are concepts with higher abstract level compared to the ones of general visual tasks (e.g., persons and cars). Therefore, classification models using only visual features often fail to achieve satisfactory performance. In this paper, a novel approach based on a "Detector-Encoder-Classifier" framework is proposed. Instead of using visual features of the whole image directly as common image-level models based on convolutional neural networks (CNNs) do, the proposed framework firstly obtains the bounding boxes of buildings in street view images from a detector. Their contextual information such as the co-occurrence patterns of building classes and their layout are then encoded into metadata by the proposed algorithm "CODING" (Context encOding of Detected buildINGs). Finally, these bounding box metadata are classified by a recurrent neural network (RNN). In addition, we made a dual-labeled dataset named "BEAUTY" (Building dEtection And Urban funcTional-zone portraYing) of 19,070 street view images and 38,857 buildings based on the existing BIC GSV [1]. The dataset can be used not only for street view image classification, but also for multi-class building detection. Experiments on "BEAUTY" show that the proposed approach achieves a 12.65% performance improvement on macro-precision and 12% on macro-recall over image-level CNN based models. Our code and dataset are available at https://github.com/kyle-one/Context-Encoding-of-Detected-Buildings/

📄 PDF Abstract BibTeX arXiv:2010.01305

Code (1)

kyle-one/Context-Encoding-of-Detected-Buildings 공식 구현 pytorch

Tasks

AllGeneral Classificationimage-classificationImage Classification

Similar Papers 제목 키워드 기반

MagicDrive: Street View Generation with Diverse 3D Geometry Control

2023-10-04 · Ruiyuan Gao, Kai Chen, Enze Xie, Lanqing Hong 외

Recent advancements in diffusion models have significantly enhanced the data synthesis with 2D control. Yet, precise 3D control in street view generation, crucial for 3D perception tasks, remains elusive. Specifically, u…

3D geometry3D Object DetectionBEV SegmentationObject+2

Leveraging Image based Prior for Visual Place Recognition

2015-05-13 · Tsukamoto Taisho, Tanaka Kanji

In this study, we propose a novel scene descriptor for visual place recognition. Unlike popular bag-of-words scene descriptors which rely on a library of vector quantized visual features, our proposed descriptor is based…

Visual Place Recognition

On Boosting Semantic Street Scene Segmentation with Weak Supervision

2019-03-08 · Panagiotis Meletis, Gijs Dubbelman

Training convolutional networks for semantic segmentation requires per-pixel ground truth labels, which are very time consuming and hence costly to obtain. Therefore, in this work, we research and develop a hierarchical …

GPUScene SegmentationSegmentationSemantic Segmentation

Fast and Regularized Reconstruction of Building Façades from Street-View Images using Binary Integer Programming

2020-02-20 · Han Hu, Libin Wang, Mier Zhang, Yulin Ding 외

Regularized arrangement of primitives on building fa\c{c}ades to aligned locations and consistent sizes is important towards structured reconstruction of urban environment. Mixed integer linear programing was used to sol…

3D Reconstruction

DiffPlace: Street View Generation via Place-Controllable Diffusion Model Enhancing Place Recognition

2026-02-12 · Ji Li, Zhiwei Li, Shihao Li, Zhenjiang Yu 외 arxiv

Generative models have advanced significantly in realistic image synthesis, with diffusion models excelling in quality and stability. Recent multi-view diffusion models improve 3D-aware street view generation, but they s…

Visual Place RecognitionContrastive LearningAutonomous DrivingImage Generation