paper-with-me

홈 › Papers

Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence

2025-10-18 · Qiongyan Wang, Xingchen Zou, Yutian Jiang, Haomin Wen, Jiaheng Wei, Qingsong Wen, Yuxuan Liang arxiv

Rapid urbanization intensifies the demand for Urban General Intelligence (UGI), referring to AI systems that can understand and reason about complex urban environments. Recent studies have built urban foundation models using supervised fine-tuning (SFT) of LLMs and MLLMs, yet these models exhibit persistent geospatial bias, producing regionally skewed predictions and limited generalization. To this end, we propose Urban-R1, a reinforcement learning-based post-training framework that aligns MLLMs with the objectives of UGI. Urban-R1 adopts Group Relative Policy Optimization (GRPO) to optimize reasoning across geographic groups and employs urban region profiling as a proxy task to provide measurable rewards from multimodal urban data. Extensive experiments across diverse regions and tasks show that Urban-R1 effectively mitigates geo-bias and improves cross-region generalization, outperforming both SFT-trained and closed-source models. Our results highlight reinforcement learning alignment as a promising pathway toward equitable and trustworthy urban intelligence.

📄 PDF Abstract BibTeX arXiv:2510.16555

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes

2025-11-27 · Di Wang, Shunyu Liu, Wentao Jiang, Fengxiang Wang 외 arxiv

Multimodal large language models (MLLMs) have undergone rapid development in advancing geospatial scene understanding. Recent studies have sought to enhance the reasoning capabilities of remote sensing MLLMs, typically t…

Reinforcement LearningScene Understanding

GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning

2025-02-25 · Theodor Lundqvist, Ludvig Delvret

Existing methods for self-supervised representation learning of geospatial regions and map entities rely extensively on the design of pretext tasks, often involving augmentations or heuristic sampling of positive and neg…

Representation Learning

Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs

2023-11-24 · Jonathan Roberts, Timo Lüddecke, Rehan Sheikh, Kai Han 외

Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains are yet to be explored, despite potenti…

Disaster Response

MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs

2024-06-24 · Wenqian Ye, Guangtao Zheng, Yunsheng Ma, Xu Cao 외

Spurious bias, a tendency to use spurious correlations between non-essential input attributes and target variables for predictions, has revealed a severe robustness pitfall in deep learning models trained on single modal…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective

2024-03-27 · Meiqi Chen, Yixin Cao, Yan Zhang, Chaochao Lu

Recent advancements in Large Language Models (LLMs) have facilitated the development of Multimodal LLMs (MLLMs). Despite their impressive capabilities, MLLMs often suffer from over-reliance on unimodal biases (e.g., lang…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)