paper-with-me

홈 › Papers

GeoMeter: Probing Depth and Height Perception of Large Visual-Language Models

2024-08-21 · Shehreen Azad, Yash Jain, Rishit Garg, Yogesh S Rawat, Vibhav Vineet

Geometric understanding is crucial for navigating and interacting with our environment. While large Vision Language Models (VLMs) demonstrate impressive capabilities, deploying them in real-world scenarios necessitates a comparable geometric understanding in visual perception. In this work, we focus on the geometric comprehension of these models; specifically targeting the depths and heights of objects within a scene. Our observations reveal that, although VLMs excel in basic geometric properties perception such as shape and size, they encounter significant challenges in reasoning about the depth and height of objects. To address this, we introduce GeoMeter, a suite of benchmark datasets encompassing Synthetic 2D, Synthetic 3D, and Real-World scenarios to rigorously evaluate these aspects. We benchmark 17 state-of-the-art VLMs using these datasets and find that they consistently struggle with both depth and height perception. Our key insights include detailed analyses of the shortcomings in depth and height reasoning capabilities of VLMs and the inherent bias present in these models. This study aims to pave the way for the development of VLMs with enhanced geometric understanding, crucial for real-world applications.

📄 PDF Abstract BibTeX arXiv:2408.11748

Code (1)

sacrcv/dh-bench 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Geometer: Graph Few-Shot Class-Incremental Learning via Prototype Representation

2022-05-27 · Bin Lu, Xiaoying Gan, Lina Yang, Weinan Zhang 외

With the tremendous expansion of graphs data, node classification shows its great importance in many real-world applications. Existing graph neural network based methods mainly focus on classifying unlabeled nodes within…

class-incremental learningClass Incremental LearningFew-Shot Class-Incremental LearningGraph Neural Network+3

BEVHeight++: Toward Robust Visual Centric 3D Object Detection

2023-09-28 · Lei Yang, Tao Tang, Jun Li, Peng Chen 외

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the percepti…

3D Object DetectionAutonomous DrivingObjectobject-detection+1

BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection

2023-03-15 · CVPR 2023 1 · Lei Yang, Kaicheng Yu, Tao Tang, Jun Li 외

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the percepti…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection

Enhancing Perception and Immersion in Pre-Captured Environments through Learning-Based Eye Height Adaptation

2023-08-24 · Qi Feng, Hubert P. H. Shum, Shigeo Morishima

Pre-captured immersive environments using omnidirectional cameras provide a wide range of virtual reality applications. Previous research has shown that manipulating the eye height in egocentric virtual environments can …

Semantic Segmentation

HeightFormer: Explicit Height Modeling without Extra Data for Camera-only 3D Object Detection in Bird's Eye View

2023-07-25 · Yiming Wu, Ruixiang Li, Zequn Qin, Xinhai Zhao 외

Vision-based Bird's Eye View (BEV) representation is an emerging perception formulation for autonomous driving. The core challenge is to construct BEV space with multi-camera features, which is a one-to-many ill-posed pr…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection