paper-with-me

Papers

ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Sensing

2025-07-07 · Zhenghui Zhao, Chen Wu, Xiangyong Cao, Di Wang, Hongruixuan Chen, Datao Tang, Liangpei Zhang, Zhuo Zheng arxiv

Spatiotemporal image generation is a highly meaningful task, which can generate future scenes conditioned on given observations. However, existing change generation methods can only handle event-driven changes (e.g., new buildings) and fail to model cross-temporal variations (e.g., seasonal shifts). In this work, we propose ChangeBridge, a conditional spatiotemporal image generation model for remote sensing. Given pre-event images and multimodal event controls, ChangeBridge generates post-event scenes that are both spatially and temporally coherent. The core idea is a drift-asynchronous diffusion bridge. Specifically, it consists of three main modules: a) Composed Bridge Initialization, which replaces noise initialization. It starts the diffusion from a composed pre-event state, modeling a diffusion bridge process. b) Asynchronous Drift Diffusion, which uses a pixel-wise drift map, assigning different drift magnitudes to event and temporal evolution. This enables differentiated generation during the pre-to-post transition. c) Drift-Aware Denoising, which embeds the drift map into the denoising network, guiding drift-aware reconstruction. Experiments show that ChangeBridge can generate better cross-spatiotemporal aligned scenarios compared to state-of-the-art methods. Additionally, ChangeBridge shows great potential for land-use planning and as a data generation engine for a series of change detection tasks. Code is available at https://github.com/zhenghuizhao/ChangeBridge

📄 PDF Abstract BibTeX arXiv:2507.04678

Code (0)

등록된 구현이 없습니다.

Tasks

Change DetectionImage Generation

Similar Papers 제목 키워드 기반

FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls

2024-04-02 · Tao Hu, Fangzhou Hong, Zhaoxi Chen, Ziwei Liu

We present FashionEngine, an interactive 3D human generation and editing system that creates 3D digital humans via user-friendly multimodal controls such as natural languages, visual perceptions, and hand-drawing sketche…

Virtual Try-on

Canvas-to-Image: Compositional Image Generation with Multimodal Controls

2025-11-26 · Yusuf Dalva, Guocheng Gordon Qian, Maya Goldenberg, Tsai-Shien Chen 외 arxiv

While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when users simultaneously specify text prompts,…

Text-to-Image GenerationSpatial Reasoning

Designing streetscapes from street-view imagery using diffusion models

2026-05-17 · Yuzhou Chen, Yuebing Liang, Lingqian Hu, Kailai Sun 외 arxiv

Street-view imagery (SVI) is widely used to quantify key indicators of urban environment, such as green- ery, sky, or road view indices. However, existing studies largely focus on measuring current streetscapes and rarel…

Scene Generation

Bridging the Gap Between Multimodal Foundation Models and World Models

2025-10-04 · Xuehai He arxiv

Humans understand the world through the integration of multiple sensory modalities, enabling them to perceive, reason about, and imagine dynamic physical processes. Inspired by this capability, multimodal foundation mode…

Causal Inference

ZeroGen: Zero-shot Multimodal Controllable Text Generation with Multiple Oracles

2023-06-29 · Haoqin Tu, Bowen Yang, Xianfeng Zhao

Automatically generating textual content with desired attributes is an ambitious task that people have pursued long. Existing works have made a series of progress in incorporating unimodal controls into language models (…

News GenerationSentenceText Generation