paper-with-me

홈 › Papers

Plan2Map: A Multimodal Benchmark for Document-Grounded Geospatial Boundary Reconstruction from Planning Records

2026-06-01 · Fabian Degen, Oishi Deb, Jindong Gu, Junchi Yu, Samuele Marro, Philip Torr, Jialin Yu arxiv

Planning records define restrictions over geographic areas, but their source documents often provide only indirect spatial evidence rather than machine-readable boundaries. We introduce Plan2Map, a 208-case multimodal benchmark for document-grounded geospatial boundary reconstruction from UK planning records. Given only a source planning document, systems must reconstruct a valid geospatial boundary from notice text, schedules, map plates, map labels, and boundary annotations; the reference GeoJSON is held out for scoring. We propose GeoPlanAgent, a document-grounded, geospatial-tool-in-the-loop system that decomposes the task into evidence extraction, localisation, map registration, boundary segmentation, projection, and verification. On Plan2Map, GeoPlanAgent achieves 0.736 mean IoU and 0.904 median IoU, with 67.8\% of predictions at or above 0.8 IoU, substantially outperforming direct VLM-to-GeoJSON baselines. Diagnostic analysis shows that direct VLM prediction remains unreliable, while remaining errors are concentrated in localisation and map registration, and supervised boundary segmentation substantially improves pixel-level mask quality. Plan2Map provides a concrete testbed for multimodal geospatial reconstruction from public planning records. Project page: https://odeb1.github.io/Plan2Map_Project_Page/.

📄 PDF Abstract BibTeX arXiv:2606.02747

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support

2026-04-14 · Muhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan 외 arxiv

Climate decision-making in the GCC states increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guidance, yet general-purpose large language models (LLMs) remain wea…

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

2026-06-11 · Dat Tien Nguyen, Thao Nguyen, Fadillah Adamsyah Maani, Huy M. Le 외 arxiv

Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate fo…

EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents

2026-05-02 · Sai Ma, Zhuang Li, Sichao Li, Xinyue Xu 외 arxiv

Earth Observation (EO) analysis is inherently interactive: resolving uncertainty often requires expanding the region of interest, retrieving historical observations, and switching across sensors such as optical and Synth…

MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Mars

2026-02-15 · Shuoyuan Wang, Yiran Wang, Hongxin Wei arxiv

Data-driven approaches like deep learning are rapidly advancing planetary science, particularly in Mars exploration. Despite recent progress, most existing benchmarks remain confined to closed-set supervised visual tasks…

Text Retrieval

MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering

2025-11-15 · Seokwon Song, Minsu Park, Gunhee Kim arxiv

Source attribution aims to enhance the reliability of AI-generated answers by including references for each statement, helping users validate the provided answers. However, existing work has primarily focused on text-onl…

Mitigating Contextual BiasVisual Question Answering