paper-with-me

홈 › Papers

ELEV-VISION-SAM: Integrated Vision Language and Foundation Model for Automated Estimation of Building Lowest Floor Elevation

2024-04-19 · Yu-Hsuan Ho, Longxiang Li, Ali Mostafavi

Street view imagery, aided by advancements in image quality and accessibility, has emerged as a valuable resource for urban analytics research. Recent studies have explored its potential for estimating lowest floor elevation (LFE), offering a scalable alternative to traditional on-site measurements, crucial for assessing properties' flood risk and damage extent. While existing methods rely on object detection, the introduction of image segmentation has broadened street view images' utility for LFE estimation, although challenges still remain in segmentation quality and capability to distinguish front doors from other doors. To address these challenges in LFE estimation, this study integrates the Segment Anything model, a segmentation foundation model, with vision language models to conduct text-prompt image segmentation on street view images for LFE estimation. By evaluating various vision language models, integration methods, and text prompts, we identify the most suitable model for street view image analytics and LFE estimation tasks, thereby improving the availability of the current LFE estimation model based on image segmentation from 33% to 56% of properties. Remarkably, our proposed method significantly enhances the availability of LFE estimation to almost all properties in which the front door is visible in the street view image. Also the findings present the first baseline and comparison of various vision models of street view image-based LFE estimation. The model and findings not only contribute to advancing street view image segmentation for urban analytics but also provide a novel approach for image segmentation tasks for other civil engineering and infrastructure analytics tasks.

📄 PDF Abstract BibTeX arXiv:2404.12606

Code (0)

등록된 구현이 없습니다.

Tasks

Image Segmentationobject-detectionObject DetectionSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

CyCLeGen: Cycle-Consistent Layout Prediction and Image Generation in Vision Foundation Models

2026-03-16 · Xiaojun Shan, Haoyu Shen, Yucheng Mao, Xiang Zhang 외 arxiv

We present CyCLeGen, a unified vision-language foundation model capable of both image understanding and image generation within a single autoregressive framework. Unlike existing vision models that depend on separate mod…

Reinforcement LearningImage Generation

Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image

2024-07-07 · Pengkun Jiao, Na Zhao, Jingjing Chen, Yu-Gang Jiang

Open-vocabulary 3D object detection (OV-3DDet) aims to localize and recognize both seen and previously unseen object categories within any new 3D scene. While language and vision foundation models have achieved success i…

3D Object DetectionObjectobject-detectionObject Detection

A Multi-Modal Foundation Model to Assist People with Blindness and Low Vision in Environmental Interaction

2023-10-31 · Yu Hao, Fan Yang, Hao Huang, Shuaihang Yuan 외

People with blindness and low vision (pBLV) encounter substantial challenges when it comes to comprehensive scene recognition and precise object identification in unfamiliar environments. Additionally, due to the vision …

Language ModelingLanguage ModellingPrompt EngineeringScene Recognition

Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models

2026-07-12 · Xiatao Sun, Yuan Zhuang, Mateo Sanchez Lopez Negrete, Matei-Victor Coldea 외 arxiv

Robotic foundation models have recently made substantial progress in multi-task capability, cross-embodiment transfer, and language-conditioned control. Yet robust deployment across diverse real-world settings remains di…

Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models

2026-03-16 · Yulin Luo, Hao Chen, Zhuangzhe Wu, Bowen Sui 외 arxiv

Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation, in which reliable action prediction critically depends on accurately interpreting and integrating visual observa…