paper-with-me

홈 › Papers

Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment

2026-06-10 · Yijie Deng, He Zhu, Wen Wang, Junyou Su, Minxin Chen, Wenjia Zhang arxiv

Problem, Research Strategy, and Findings: The rise of large language models (LLMs) raises a key question for urban planning: which forms of professional planning knowledge can AI replicate, and which still require human judgment? Although AI tools are increasingly used in planning practice, there is still no systematic framework for testing whether they can reason with the contextual sensitivity, value awareness, and institutional literacy central to planning expertise. This paper introduces Urban Planning Bench (UPBench), a domain-specific evaluation framework that assesses LLM reasoning through a 4x5 matrix of four knowledge pillars and five cognitive levels adapted from Bloom's revised taxonomy. Evaluating 25 LLMs with automated scoring and expert review, we find a non-monotonic cognitive curve: models perform better on higher-order analytical tasks than on factual recall and integrative judgment. This suggests that planning knowledge often treated as lower-order is deeply shaped by institutional, jurisdictional, and temporal context, making it hard for LLMs to generalize. We summarize these limits as four epistemic diagnostics: regulatory hallucination, conceptual conflation, wickedness paralysis, and phronetic deficit. Takeaway for Practice: The findings support differential delegation in planning. LLMs can assist with cross-disciplinary synthesis, literature review, scenario generation, and preliminary policy analysis. However, they remain unreliable for jurisdiction-specific regulation, normative conflict resolution, and context-sensitive procedure. Agencies should require verification for AI-assisted regulatory analysis, while planning education should emphasize institutional literacy, normative judgment, and contextual sensitivity.

📄 PDF Abstract BibTeX arXiv:2606.11678

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

2025-03-08 · Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang 외

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a benchmark to evaluate whether video-large l…

BenchmarkingcounterfactualMultiple-choice

Towards Urban Planing AI Agent in the Age of Agentic AI

2025-07-19 · Rui Liu, Tao Zhe, Zhong-Ren Peng, Necati Catbas 외 arxiv

Generative AI, large language models, and agentic AI have emerged separately of urban planning. However, the convergence between AI and urban planning presents an interesting opportunity towards AI urban planners. Existi…

BAPPA: Benchmarking Agents, Plans, and Pipelines for Automated Text-to-SQL Generation

2025-11-06 · Fahim Ahmed, Md Mubtasim Ahasan, Jahir Sadik Monon, Muntasir Wahed 외 arxiv

Text-to-SQL systems provide a natural language interface that can enable even laymen to access information stored in databases. However, existing Large Language Models (LLM) struggle with SQL generation from natural inst…

PlanGPT: Enhancing Urban Planning with Tailored Language Model and Efficient Retrieval

2024-02-29 · He Zhu, Wenjia Zhang, Nuoxian Huang, Boyang Li 외

In the field of urban planning, general-purpose large language models often struggle to meet the specific needs of planners. Tasks like generating urban planning texts, retrieving related information, and evaluating plan…

Language ModelingLanguage ModellingLarge Language ModelRetrieval

HYPE: Hybrid Planning with Ego Proposal-Conditioned Predictions

2025-10-14 · Hang Yu, Julian Jordan, Julian Schmidt, Silvan Lindner 외 arxiv

Safe and interpretable motion planning in complex urban environments needs to reason about bidirectional multi-agent interactions. This reasoning requires to estimate the costs of potential ego driving maneuvers. Many ex…

Motion Planning