paper-with-me

홈 › Papers

Joint Multi-Feature Spatial Context for Scene Recognition on the Semantic Manifold

2015-06-01 · CVPR 2015 6 · Xinhang Song, Shuqiang Jiang, Luis Herranz

In the semantic multinomial framework patches and images are modeled as points in a semantic probability simplex. Patch theme models are learned resorting to weak supervision via image labels, which leads the problem of scene categories co-occurring in this semantic space. Fortunately, each category has its own co-occurrence patterns that are consistent across the images in that category. Thus, discovering and modeling these patterns is critical to improve the recognition performance in this representation. In this paper, we observe that not only global co-occurrences at the image-level are important, but also different regions have different category co-occurrence patterns. We exploit local contextual relations to address the problem of discovering consistent co-occurrence patterns and removing noisy ones. Our hypothesis is that a less noisy semantic representation, would greatly help the classifier to model consistent co-occurrences and discriminate better between scene categories. An important advantage of modeling features in a semantic space, is that this space is feature independent. Thus, we can combine multiple features and spatial neighbors in the same common space, and formulate the problem as minimizing a context-dependent energy. Experimental results show that exploiting different types of contextual relations consistently improves the recognition accuracy. In particular, larger datasets benefit more from the proposed method, leading to very competitive performance.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Recognition

Similar Papers 제목 키워드 기반

Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks

2026-02-14 · Guanfeng Tang, Hongbo Zhao, Ziwei Long, Jiayao Li 외 arxiv

Inspired by the human visual system, which operates on two parallel yet interactive streams for contextual and spatial understanding, this article presents Two Interactive Streams (TwInS), a novel bio-inspired joint lear…

Scene Parsing

Pose-Guided Temporal Enhancement for Robust Low-Resolution Hand Reconstruction

2025-01-01 · CVPR 2025 1 · Kaixin Fan, Pengfei Ren, Jingyu Wang, Haifeng Sun 외

3D hand reconstruction is essential in non-contact human-computer interaction applications, but existing methods struggle with low-resolution images, which occur in slightly distant interactive scenes. Leveraging tem…

Geospatial-Prior Guidance for 3D Semantic Scene Completion

2026-08-04 · Meng Wang, Shougao Zhang, Wenzhe He, Ruihui Li 외 arxiv

Inferring complete 3D geometry and semantics from onboard images remains challenging because occlusions and restricted fields of view leave large scene regions underconstrained. Although satellite imagery provides wide-a…

3D Semantic Scene Completion

Feature boosting with efficient attention for scene parsing

2024-02-29 · Vivek Singh, Shailza Sharma, Fabio Cuzzolin

The complexity of scene parsing grows with the number of object and scene classes, which is higher in unrestricted open scenes. The biggest challenge is to model the spatial relation between scene elements while succeedi…

Scene Parsing

JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling

2024-08-07 · Seok Hwan Lee, Taein Son, Soo Won Seo, Jisong Kim 외

Video action detection (VAD) is a formidable vision task that involves the localization and classification of actions within the spatial and temporal dimensions of a video clip. Among the myriad VAD architectures, two-st…

Action DetectionRelationVideo Action Detection