paper-with-me

홈 › Papers

SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis

2024-03-14 · Huan-ang Gao, Mingju Gao, Jiaju Li, Wenyi Li, Rong Zhi, Hao Tang, Hao Zhao

Semantic image synthesis (SIS) shows good promises for sensor simulation. However, current best practices in this field, based on GANs, have not yet reached the desired level of quality. As latent diffusion models make significant strides in image generation, we are prompted to evaluate ControlNet, a notable method for its dense control capabilities. Our investigation uncovered two primary issues with its results: the presence of weird sub-structures within large semantic areas and the misalignment of content with the semantic mask. Through empirical study, we pinpointed the cause of these problems as a mismatch between the noised training data distribution and the standard normal prior applied at the inference stage. To address this challenge, we developed specific noise priors for SIS, encompassing spatial, categorical, and a novel spatial-categorical joint prior for inference. This approach, which we have named SCP-Diff, has set new state-of-the-art results in SIS on Cityscapes, ADE20K and COCO-Stuff, yielding a FID as low as 10.53 on Cityscapes. The code and models can be accessed via the project page.

📄 PDF Abstract BibTeX arXiv:2403.09638

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Mixed Continuous and Categorical Flow Matching for 3D De Novo Molecule Generation

2024-04-30 · Ian Dunn, David Ryan Koes

Deep generative models that produce novel molecular structures have the potential to facilitate chemical discovery. Diffusion models currently achieve state of the art performance for 3D molecule generation. In this work…

3D Molecule Generation

TabDLM: Free-Form Tabular Data Generation via Joint Numerical-Language Diffusion

2026-02-26 · Donghong Cai, Jiarui Feng, Yanbo Wang, Da Zheng 외 arxiv

Synthetic tabular data generation has attracted growing attention due to its importance for data augmentation, foundation models, and privacy. However, real-world tabular datasets increasingly contain free-form text fiel…

Tabular Data GenerationData Augmentation

Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser

2024-03-07 · Qingyuan Cai, Xuecai Hu, Saihui Hou, Li Yao 외

Recently, diffusion-based methods for monocular 3D human pose estimation have achieved state-of-the-art (SOTA) performance by directly regressing the 3D joint coordinates from the 2D pose sequence. Although some methods …

3D Human Pose EstimationDisentanglementMonocular 3D Human Pose EstimationMulti-Hypotheses 3D Human Pose Estimation+1

Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics

2026-06-10 · Sam Gijsen, Michał Łukomski, Marc-André Schulz, Kerstin Ritter arxiv

Flow matching and diffusion models enable conditional generation across domains ranging from images to proteins, with recent extensions to out-of-distribution contexts. Yet generative models of neural time series have la…

Zero-shot Generalization

SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis

2025-08-26 · Xiaohao Sun, Divyam Goel, Angel X. Chang arxiv

We present SemLayoutDiff, a unified model for synthesizing diverse 3D indoor scenes across multiple room types. The model introduces a scene layout representation combining a top-down semantic map and attributes for each…

Indoor Scene Synthesis