paper-with-me

홈 › Papers

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View

2026-06-01 · Dongdong Wang, Alina Hagen, Isabelle Gatmaitan, Hao Zhou, Yiwen Dong, Shabboo Valipoor, Vivian W. H. Wong, Lingyao Li arxiv

Assessing built-environment interaction, such as wheelchair accessibility, is difficult because real-world mobility is shaped by distributed, context-dependent, and temporary barriers that are hard to capture at scale. To support scalable assessment, this paper examines whether vision-language models (VLMs) can identify accessibility barriers from Google Street View (GSV) imagery. We propose an expert-guided retrieval-augmented framework that combines GSV images, ADA-informed guidance, and expert-derived rubrics to evaluate accessibility dimensions. We collect a campus-scale dataset at the University of Florida, linking 407 unique GSV locations with GPS-derived wheelchair dwell behavior as a mobility-friction signal. Results show that VLM ratings are both negatively correlated and distributionally similar with dwell time, indicating partial but consistent alignment with a behavioral proxy for mobility friction. Visual cue analysis shows that certain environmental objects, such as curb ramps and crosswalks, are associated with higher VLM accessibility scores, while alignment remains limited for subtle surface conditions, transient obstructions, and viewpoint-dependent barriers. Overall, our findings show the potential of expert-guided VLMs for scalable accessibility assessment aligning with sensor-derived indicators of real-world wheelchair navigation.

📄 PDF Abstract BibTeX arXiv:2606.07642

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What data do we need for training an AV motion planner?

2021-05-26 · Long Chen, Lukas Platinsky, Stefanie Speichert, Blazej Osinski 외

We investigate what grade of sensor data is required for training an imitation-learning-based AV planner on human expert demonstration. Machine-learned planners are very hungry for training data, which is usually collect…

Imitation LearningMotion Planning

Feel the Force: Contact-Driven Learning from Humans

2025-06-02 · Ademi Adeniji, Zhuoran Chen, Vincent Liu, Venkatesh Pattabiraman 외

Controlling fine-grained forces during manipulation remains a core challenge in robotics. While robot policies learned from robot-collected data or simulation show promise, they struggle to generalize across the diverse …

What is the people posting about symptoms related to Coronavirus in Bogota, Colombia?

2020-03-25 · Josimar E. Chire Saire, Roberto C. Navarro

During the last months, there is an increasing alarm about a new mutation of coronavirus, covid-19 coined by World Health Organization(WHO) with an impact in many areas: economy, health, politics and others. This situati…

VLM3: Vision Language Models Are Native 3D Learners

2026-05-28 · Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu 외 arxiv

Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still largely relies on exp…

Camera Pose EstimationDepth Estimation

AIDE: Agentically Improve Visual Language Model with Domain Experts

2025-02-13 · Ming-Chang Chiu, Fuxiao Liu, Karan Sapra, Andrew Tao 외

The enhancement of Visual Language Models (VLMs) has traditionally relied on knowledge distillation from larger, more capable models. This dependence creates a fundamental bottleneck for improving state-of-the-art system…

Knowledge DistillationLanguage ModelingLanguage ModellingMME