paper-with-me

홈 › Papers

ASK: Adaptively Selecting Key Local Features for RGB-D Scene Recognition

2021-10-14 · Zhitong Xiong, Yuan Yuan, Qi Wang

Indoor scene images usually contain scattered objects and various scene layouts, which make RGB-D scene classification a challenging task. Existing methods still have limitations for classifying scene images with great spatial variability. Thus, how to extract local patch-level features effectively using only image labels is still an open problem for RGB-D scene recognition. In this paper, we propose an efficient framework for RGB-D scene recognition, which adaptively selects important local features to capture the great spatial variability of scene images. Specifically, we design a differentiable local feature selection (DLFS) module, which can extract the appropriate number of key local scenerelated features. Discriminative local theme-level and object-level representations can be selected with the DLFS module from the spatially-correlated multi-modal RGB-D features. We take advantage of the correlation between RGB and depth modalities to provide more cues for selecting local features. To ensure that discriminative local features are selected, the variational mutual information maximization loss is proposed. Additionally, the DLFS module can be easily extended to select local features of different scales. By concatenating the local-orderless and global structured multi-modal features, the proposed framework can achieve state-of-the-art performance on public RGB-D scene recognition datasets.

📄 PDF Abstract BibTeX arXiv:2110.07703

Code (0)

등록된 구현이 없습니다.

Tasks

feature selectionScene ClassificationScene Recognition

Methods 이 논문이 사용한 방법론

Feature Selection Feature selection, also known as variable selection, attribute selection or variable subset selection, is the process of selecting a subset of relevant features (variables,…

Similar Papers 제목 키워드 기반

Visual Concept Reasoning Networks

2020-08-26 · Taesup Kim, Sungwoong Kim, Yoshua Bengio

A split-transform-merge strategy has been broadly used as an architectural constraint in convolutional neural networks for visual recognition tasks. It approximates sparsely connected networks by explicitly defining mult…

Action Recognitionimage-classificationImage Classificationobject-detection+3

Dynamic Aggregated Network for Gait Recognition

2023-01-01 · CVPR 2023 1 · Kang Ma, Ying Fu, Dezhi Zheng, Chunshui Cao 외

Gait recognition is beneficial for a variety of applications, including video surveillance, crime scene investigation, and social security, to mention a few. However, gait recognition often suffers from multiple exte…

Gait Recognition

Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance

2024-12-13 · Jiahao Lyu, Wei Wang, Dongbao Yang, Jinwen Zhong 외

Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determ…

Scene Text RecognitionText Spotting

Hierarchical Fusion of Local and Global Visual Features with Mixture-of-Experts for Remote Sensing Image Scene Classification

2025-10-31 · Yuanhao Tang, Xuechao Zou, Zhengpei Hu, Junliang Xing 외 arxiv

Remote sensing image scene classification remains a challenging task, primarily due to the complex spatial structures and multi-scale characteristics of ground objects. Although CNN-based methods excel at extracting loca…

Scene ClassificationScene Recognition

DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition

2025-08-07 · Haijing Liu, Tao Pu, Hefeng Wu, Keze Wang 외 arxiv

Open-Vocabulary Multi-Label Recognition (OV-MLR) aims to identify multiple seen and unseen object categories within an image, requiring both precise intra-class localization to pinpoint objects and effective inter-class …