paper-with-me

Papers

Harvesting Discriminative Meta Objects with Deep CNN Features for Scene Classification

2015-10-06 · ICCV 2015 12 · Ruobing Wu, Baoyuan Wang, Wenping Wang, Yizhou Yu

Recent work on scene classification still makes use of generic CNN features in a rudimentary manner. In this ICCV 2015 paper, we present a novel pipeline built upon deep CNN features to harvest discriminative visual objects and parts for scene classification. We first use a region proposal technique to generate a set of high-quality patches potentially containing objects, and apply a pre-trained CNN to extract generic deep features from these patches. Then we perform both unsupervised and weakly supervised learning to screen these patches and discover discriminative ones representing category-specific objects and parts. We further apply discriminative clustering enhanced with local CNN fine-tuning to aggregate similar objects and parts into groups, called meta objects. A scene image representation is constructed by pooling the feature response maps of all the learned meta objects at multiple spatial scales. We have confirmed that the scene image representation obtained using this new pipeline is capable of delivering state-of-the-art performance on two popular scene benchmark datasets, MIT Indoor 67~\cite{MITIndoor67} and Sun397~\cite{Sun397}

📄 PDF Abstract BibTeX arXiv:1510.01440

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringGeneral ClassificationRegion ProposalScene ClassificationWeakly-supervised Learning

Similar Papers 제목 키워드 기반

Inter-object Discriminative Graph Modeling for Indoor Scene Recognition

2023-11-10 · Chuanxin Song, Hanbo Wu, Xin Ma

Variable scene layouts and coexisting objects across scenes make indoor scene recognition still a challenging task. Leveraging object information within scenes to enhance the distinguishability of feature representations…

ObjectScene Recognition

Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring

2025-11-24 · Siyuan Wei, Chunjie Wang, Xiao Liu, Xiaosheng Yan 외 arxiv

3D Multi-modal Large Language Models (MLLMs) still lag behind their 2D peers, largely because large-scale, high-quality 3D scene-dialogue datasets remain scarce. Prior efforts hinge on expensive human annotation and leav…

Visual Grounding

Exploiting Polarized Material Cues for Robust Car Detection

2024-01-05 · Wen Dong, Haiyang Mei, Ziqi Wei, Ao Jin 외

Car detection is an important task that serves as a crucial prerequisite for many automated driving functions. The large variations in lighting/weather conditions and vehicle densities of the scenes pose significant chal…

Semantic-guided modeling of spatial relation and object co-occurrence for indoor scene recognition

2023-05-22 · Chuanxin Song, Hanbo Wu, Xin Ma

Exploring the semantic context in scene images is essential for indoor scene recognition. However, due to the diverse intra-class spatial layouts and the coexisting inter-class objects, modeling contextual relationships …

RelationScene RecognitionSemantic Segmentation

KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM

2025-12-01 · Zaid Nasser, Mikhail Iumanov, Tianhao Li, Maxim Popov 외 arxiv

We present KM-ViPE (Knowledge Mapping Video Pose Engine), a real-time open-vocabulary SLAM framework for uncalibrated monocular cameras in dynamic environments. Unlike systems requiring depth sensors and offline calibrat…

Semantic SLAM