paper-with-me

홈 › Papers

From Pixels to Predicates Structuring urban perception with scene graphs

2025-12-22 · Yunlong Liu, Shuyang Li, Pengyuan Liu, Yu Zhang, Rudi Stouffs arxiv

Perception research is increasingly modelled using streetscapes, yet many approaches still rely on pixel features or object co-occurrence statistics, overlooking the explicit relations that shape human perception. This study proposes a three stage pipeline that transforms street view imagery (SVI) into structured representations for predicting six perceptual indicators. In the first stage, each image is parsed using an open-set Panoptic Scene Graph model (OpenPSG) to extract object predicate object triplets. In the second stage, compact scene-level embeddings are learned through a heterogeneous graph autoencoder (GraphMAE). In the third stage, a neural network predicts perception scores from these embeddings. We evaluate the proposed approach against image-only baselines in terms of accuracy, precision, and cross-city generalization. Results indicate that (i) our approach improves perception prediction accuracy by an average of 26% over baseline models, and (ii) maintains strong generalization performance in cross-city prediction tasks. Additionally, the structured representation clarifies which relational patterns contribute to lower perception scores in urban scenes, such as graffiti on wall and car parked on sidewalk. Overall, this study demonstrates that graph-based structure provides expressive, generalizable, and interpretable signals for modelling urban perception, advancing human-centric and context-aware urban analytics.

📄 PDF Abstract BibTeX arXiv:2512.19221

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modeling Subjective Urban Perception with Human Gaze

2026-05-01 · Lin Che, Xi Wang, Marc Pollefeys, Konrad Schindler 외 arxiv

Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computational approaches primarily model urban perception directly from street v…

Scene Understanding

UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective

2025-09-26 · Jun He, Yi Lin, Zilong Huang, Jiacong Yin 외 arxiv

Urban development impacts over half of the global population, making human-centered understanding of its structural and perceptual changes essential for sustainable development. While Multimodal Large Language Models (ML…

Scene UnderstandingChange Detection

Copy-Pasting Coherent Depth Regions Improves Contrastive Learning for Urban-Scene Segmentation

2022-11-25 · Liang Zeng, Attila Lengyel, Nergis Tömen, Jan van Gemert

In this work, we leverage estimated depth to boost self-supervised contrastive learning for segmentation of urban scenes, where unlabeled videos are readily available for training self-supervised depth estimation. We arg…

Contrastive LearningDepth EstimationScene SegmentationSegmentation+2

Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception

2026-05-11 · Yiwei Ou, Chung Ching Cheung, Jun Yang Ang, Xiaobin Ren 외 arxiv

We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imagery. The corpus contains over 2 Million public social media images and…

Cross-Modal RetrievalInstance SegmentationScene ClassificationImage Classification

TMBuD: A dataset for urban scene building detection

2021-10-27 · Orhei Ciprian, Vert Silviu, Mocofan Muguras, Vasiu Radu

Building recognition and 3D reconstruction of human made structures in urban scenarios has become an interesting and actual topic in the image processing domain. For this research topic the Computer Vision and Augmented …

3D ReconstructionSemantic Segmentation