paper-with-me

홈 › Papers

ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge

2023-11-24 · Eslam Abdelrahman, Liangbing Zhao, Vincent Tao Hu, Matthieu Cord, Patrick Perez, Mohamed Elhoseiny

Diffusion models break down the challenging task of generating data from high-dimensional distributions into a series of easier denoising steps. Inspired by this paradigm, we propose a novel approach that extends the diffusion framework into modality space, decomposing the complex task of RGB image generation into simpler, interpretable stages. Our method, termed ToddlerDiffusion, cascades modality-specific models, each responsible for generating an intermediate representation, such as contours, palettes, and detailed textures, ultimately culminating in a high-quality RGB image. Instead of relying on the naive LDM concatenation conditioning mechanism to connect the different stages together, we employ Schr\"odinger Bridge to determine the optimal transport between different modalities. Although employing a cascaded pipeline introduces more stages, which could lead to a more complex architecture, each stage is meticulously formulated for efficiency and accuracy, surpassing Stable-Diffusion (LDM) performance. Modality composition not only enhances overall performance but enables emerging proprieties such as consistent editing, interaction capabilities, high-level interpretability, and faster convergence and sampling rate. Extensive experiments on diverse datasets, including LSUN-Churches, ImageNet, CelebHQ, and LAION-Art, demonstrate the efficacy of our approach, consistently outperforming state-of-the-art methods. For instance, ToddlerDiffusion achieves notable efficiency, matching LDM performance on LSUN-Churches while operating 2$\times$ faster with a 3$\times$ smaller architecture. The project website is available at: https://toddlerdiffusion.github.io/website/

📄 PDF Abstract BibTeX arXiv:2311.14542

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingImage Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation

2024-11-12 · Yushi Lan, Shangchen Zhou, Zhaoyang Lyu, Fangzhou Hong 외

While 3D content generation has advanced significantly, existing methods still face challenges with input formats, latent space design, and output representations. This paper introduces a novel 3D generation framework th…

3D GenerationDisentanglement

Real-time interactive magnetic resonance (MR) temperature imaging in both aqueous and adipose tissues using cascaded deep neural networks for MR-guided focused ultrasound surgery (MRgFUS)

2019-08-29 · Jong-Min Kim, You-Jin Jeong, Han-Jae Chung, Chulhyun Lee 외

Purpose: To acquire the real-time interactive temperature map for aqueous and adipose tissue, the problems of long acquisition and processing time must be addressed. To overcome these major challenges, this paper propose…

Image Reconstruction

Interactive Segmentation and Report Generation for CT Images

2025-03-05 · Yannian Gu, Wenhui Lei, HanYu Chen, Xiaofan Zhang 외

Automated CT report generation plays a crucial role in improving diagnostic accuracy and clinical workflow efficiency. However, existing methods lack interpretability and impede patient-clinician understanding, while the…

AttributeDiagnosticInteractive SegmentationLesion Segmentation+1

GraphVid: Interactive Graph-Controllable Video Generation

2026-07-23 · Vedant Shah, Onkar Susladkar, Tushar Prakash, Kiet Nguyen 외 arxiv

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, tr…

Video Generation

Cascaded Diffusion Models for High Fidelity Image Generation

2021-05-30 · Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet 외

We show that cascaded diffusion models are capable of generating high fidelity images on the class-conditional ImageNet generation benchmark, without any assistance from auxiliary image classifiers to boost sample qualit…

Data AugmentationImage GenerationSuper-ResolutionVocal Bursts Intensity Prediction