paper-with-me

Papers

RoBus: A Multimodal Dataset for Controllable Road Networks and Building Layouts Generation

2024-07-10 · Tao Li, Ruihang Li, Huangnan Zheng, Shanding Ye, Shijian Li, Zhijie Pan

Automated 3D city generation, focusing on road networks and building layouts, is in high demand for applications in urban design, multimedia games and autonomous driving simulations. The surge of generative AI facilitates designing city layouts based on deep learning models. However, the lack of high-quality datasets and benchmarks hinders the progress of these data-driven methods in generating road networks and building layouts. Furthermore, few studies consider urban characteristics, which generally take graphics as analysis objects and are crucial for practical applications, to control the generative process. To alleviate these problems, we introduce a multimodal dataset with accompanying evaluation metrics for controllable generation of Road networks and Building layouts (RoBus), which is the first and largest open-source dataset in city generation so far. RoBus dataset is formatted as images, graphics and texts, with $72,400$ paired samples that cover around $80,000km^2$ globally. We analyze the RoBus dataset statistically and validate the effectiveness against existing road networks and building layouts generation methods. Additionally, we design new baselines that incorporate urban characteristics, such as road orientation and building density, in the process of generating road networks and building layouts using the RoBus dataset, enhancing the practicality of automated urban design. The RoBus dataset and related codes are published at https://github.com/tourlics/RoBus_Dataset.

📄 PDF Abstract BibTeX arXiv:2407.07835

Code (1)

tourlics/robus_dataset 공식 구현

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

2025-01-21 · Shiyue Zhang, Zheng Chong, Xi Lu, Wenqing Zhang 외

Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a promising technique, offering the potential …

Image Generation

HeartBeat: Towards Controllable Echocardiography Video Synthesis with Multimodal Conditions-Guided Diffusion Models

2024-06-20 · Xinrui Zhou, Yuhao Huang, Wufeng Xue, Haoran Dou 외

Echocardiography (ECHO) video is widely used for cardiac examination. In clinical, this procedure heavily relies on operator experience, which needs years of training and maybe the assistance of deep learning-based syste…

MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

2025-06-05 · Zikui Cai, Andrew Wang, Anirudh Satheesh, Ankit Nakhawa 외

Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static images, failing to capture the temporal c…

Dataset GenerationMathematical Problem-SolvingMultimodal Reasoning

CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion

2025-11-26 · Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu 외 arxiv

We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g., depth, edges) are insufficient: they s…

Video Generation

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

2026-05-05 · Lin Song, Wenbo Li, Guoqing Ma, Wei Tang 외 arxiv

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language M…

Text-to-Image GenerationImage Editing