paper-with-me

홈 › Papers

3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding

2026-03-24 · Yiping Chen, Jinpeng Li, Wenyu Ke, Yang Luo, Jie Ouyang, Zhongjie He, Li Liu, Hongchao Fan, Hao Wu arxiv

While multi-modality large language models excel in object-centric or indoor scenarios, scaling them to 3D city-scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity-LLM, a unified framework designed for 3D city-scale vision-language perception and understanding. 3DCity-LLM employs a coarse-to-fine feature encoding strategy comprising three parallel branches for target object, inter-object relationship, and global scene. To facilitate large-scale training, we introduce 3DCity-LLM-1.2M dataset that comprises approximately 1.2 million high-quality samples across seven representative task categories, ranging from fine-grained object analysis to multi-faceted scene planning. This strictly quality-controlled dataset integrates explicit 3D numerical information and diverse user-oriented simulations, enriching the question-answering diversity and realism of urban scenarios. Furthermore, we apply a multi-dimensional protocol based on text-similarity metrics and LLM-based semantic assessment to ensure faithful and comprehensive evaluations for all methods. Extensive experiments on two benchmarks demonstrate that 3DCity-LLM significantly outperforms existing state-of-the-art methods, offering a promising and meaningful direction for advancing spatial reasoning and urban intelligence. The source code and dataset are available at https://github.com/SYSU-3DSTAILab/3D-City-LLM.

📄 PDF Abstract BibTeX arXiv:2603.23447

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Integrating 3D City Data through Knowledge Graphs

2023-10-17 · Linfang Ding, Guohui Xiao, Albulen Pano, Mattia Fumagalli 외

CityGML is a widely adopted standard by the Open Geospatial Consortium (OGC) for representing and exchanging 3D city models. The representation of semantic and topological properties in CityGML makes it possible to query…

Knowledge Graphs

AQPDCITY Dataset: Picture-Based PM Monitoring in the Urban Area of Big Cities

2020-03-22 · Yonghui Zhang, Ke Gu

Since Particulate Matters (PMs) are closely related to people's living and health, it has become one of the most important indicator of air quality monitoring around the world. But the existing sensor-based methods for P…

WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence

2026-07-07 · Xiangyu Han, Mengyu Yang, Jiaqi Li, Bowen Chang 외 arxiv

Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI build spatial representations at a comparable scale? Although recent foundation models ha…

SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

2023-05-18 · Dong Zhang, ShiMin Li, Xin Zhang, Jun Zhan 외

Multi-modal large language models are regarded as a crucial step towards Artificial General Intelligence (AGI) and have garnered significant interest with the emergence of ChatGPT. However, current speech-language models…

Language ModelingLanguage ModellingLarge Language ModelSpeech Recognition+1

LMFusion: Adapting Pretrained Language Models for Multimodal Generation

2024-12-19 · Weijia Shi, Xiaochuang Han, Chunting Zhou, Weixin Liang 외

We present LMFusion, a framework for empowering pretrained text-only large language models (LLMs) with multimodal generative capabilities, enabling them to understand and generate both text and images in arbitrary sequen…

Image Generationmultimodal generation