paper-with-me

Papers

Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection

2025-07-17 · Jingyao Wang, Yiming Chen, Lingyu Si, Changwen Zheng

Scene understanding is one of the core tasks in computer vision, aiming to extract semantic information from images to identify objects, scene categories, and their interrelationships. Although advancements in Vision-Language Models (VLMs) have driven progress in this field, existing VLMs still face challenges in adaptation to unseen complex wide-area scenes. To address the challenges, this paper proposes a Hierarchical Coresets Selection (HCS) mechanism to advance the adaptation of VLMs in complex wide-area scene understanding. It progressively refines the selected regions based on the proposed theoretically guaranteed importance function, which considers utility, representativeness, robustness, and synergy. Without requiring additional fine-tuning, HCS enables VLMs to achieve rapid understandings of unseen scenes at any scale using minimal interpretable regions while mitigating insufficient feature density. HCS is a plug-and-play method that is compatible with any VLM. Experiments demonstrate that HCS achieves superior performance and universality in various tasks.

📄 PDF Abstract BibTeX arXiv:2507.13061

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

MM-3DScene: 3D Scene Understanding by Customizing Masked Modeling with Informative-Preserved Reconstruction and Self-Distilled Consistency

2022-12-20 · CVPR 2023 1 · Mingye Xu, Mutian Xu, Tong He, Wanli Ouyang 외

Masked Modeling (MM) has demonstrated widespread success in various vision challenges, by reconstructing masked visual patches. Yet, applying MM for large-scale 3D scenes remains an open problem due to the data sparsity …

object-detectionObject DetectionScene UnderstandingSemantic Segmentation

OpenSU3D: Open World 3D Scene Understanding using Foundation Models

2024-07-19 · Rafay Mohiuddin, Sai Manoj Prakhya, Fiona Collins, Ziyuan Liu 외

In this paper, we present a novel, scalable approach for constructing open set, instance-level 3D scene representations, advancing open world understanding of 3D environments. Existing methods require pre-constructed 3D …

Scene UnderstandingSpatial ReasoningZero-shot Generalization

Advancing the Understanding of Fine-Grained 3D Forest Structures using Digital Cousins and Simulation-to-Reality: Methods and Datasets

2025-01-07 · Jing Liu, Duanchu Wang, Haoran Gong, Chongyu Wang 외

Understanding and analyzing the spatial semantics and structure of forests is essential for accurate forest resource monitoring and ecosystem research. However, the lack of large-scale and annotated datasets has limited …

Data Augmentationparameter estimationScene UnderstandingSynthetic Data Generation

Scene Understanding Networks for Autonomous Driving based on Around View Monitoring System

2018-05-18 · JeongYeol Baek, Ioana Veronica Chelu, Livia Iordache, Vlad Paunescu 외

Modern driver assistance systems rely on a wide range of sensors (RADAR, LIDAR, ultrasound and cameras) for scene understanding and prediction. These sensors are typically used for detecting traffic participants and scen…

3D Object DetectionAutonomous DrivingDrivable Area Detectionobject-detection+2

Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble

2025-01-21 · Lin Duan, Yanming Xiu, Maria Gorlatova

Augmented Reality (AR) enhances the real world by integrating virtual content, yet ensuring the quality, usability, and safety of AR experiences presents significant challenges. Could Vision-Language Models (VLMs) offer …