paper-with-me

홈 › Papers

AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery

2026-07-16 · Zahratu Shabrina, Muhammad Asa, Jin Rui, Lu Yin, Stephen Law arxiv

This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images. Predictions generated by VLMs are compared with inference by human experts (civil engineers and architects) as a source of manually labelled ground-truth data. We evaluate several state-of-the-art VLMs, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. By applying different scaling strategies and prompting techniques, we found that Chain-of-Thought prompts provide an overall more stable model performance. We also investigate the reasoning behind VLMs' building-typology predictions by examining the probabilities of keywords appearing in AI explanations. This enabled us to analyse patterns in these reasonings and identify key themes driving both agreements and disagreements between VLM and expert labels. We find that AI tends to focus on visual indicators, whereas human experts place greater emphasis on broader contextual cues and domain knowledge, in addition to visual cues. Overall, VLM can approximate experts' capability in building-typology classification at scale, with an average accuracy of approximately 70%. The study demonstrates the VLM's potential for AI automation in tasks that require pattern recognition and object identification in an urban context. AI have the potential to serve as complementary and collaborative tools for urban analysis, leveraging their strengths in understanding visual patterns. This study contributes to the exploration of the efficiency and scalability of AI visual prediction and provides insights into the reasoning processes that could support automation processes in urban analysis and prediction.

📄 PDF Abstract BibTeX arXiv:2607.14756

Code (1)

Tavish9/awesome-daily-AI-arxiv ★ 111

Similar Papers 제목 키워드 기반

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment

2025-08-07 · Chang Hong, Minghao Wu, Qingying Xiao, Yuchi Wang 외 arxiv

As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical. While achieving high accuracy on knowledge benchmarks, LLMs lack validated assessment for navigating ethic…

Steps Towards Programs that Manage Uncertainty

2013-03-27 · Paul Cohen

Reasoning under uncertainty in Al hats come to mean assessing the credibility of hypotheses inferred from evidence. But techniques for assessing credibility do not tell a problem solver what to do when it is uncertain. T…

Diagnostic

Direct Uncertainty Prediction for Medical Second Opinions

2018-07-04 · Maithra Raghu, Katy Blumer, Rory Sayres, Ziad Obermeyer 외

The issue of disagreements amongst human experts is a ubiquitous one in both machine learning and medicine. In medicine, this often corresponds to doctor disagreements on a patient diagnosis. In this work, we show that m…

BIG-bench Machine LearningGeneral ClassificationPrediction

From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome Classification

2023-10-18 · Shanshan Xu, T. Y. S. S Santosh, Oana Ichim, Isabella Risini 외

In legal NLP, Case Outcome Classification (COC) must not only be accurate but also trustworthy and explainable. Existing work in explainable COC has been limited to annotations by a single expert. However, it is well-kno…

International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty

2025-03-18 · Rebecca Scholefield, Samuel Martin, Otto Barten

The malicious use or malfunction of advanced general-purpose AI (GPAI) poses risks that, according to leading experts, could lead to the 'marginalisation or extinction of humanity.' To address these risks, there are an i…