S2ST: Image-to-Image Translation in the Seed Space of Latent Diffusion
Image-to-image translation (I2IT) refers to the process of transforming images from a source domain to a target domain while maintaining a fundamental connection in terms of image content. In the past few years, remarkable advancements in I2IT were achieved by Generative Adversarial Networks (GANs), which nevertheless struggle with translations requiring high precision. Recently, Diffusion Models have established themselves as the engine of choice for image generation. In this paper we introduce S2ST, a novel framework designed to accomplish global I2IT in complex photorealistic images, such as day-to-night or clear-to-rain translations of automotive scenes. S2ST operates within the seed space of a Latent Diffusion Model, thereby leveraging the powerful image priors learned by the latter. We show that S2ST surpasses state-of-the-art GAN-based I2IT methods, as well as diffusion-based approaches, for complex automotive scenes, improving fidelity while respecting the target domain's appearance across a variety of domains. Notably, S2ST obviates the necessity for training domain-specific translation networks.
Code (0)
등록된 구현이 없습니다.
Tasks
Image GenerationImage-to-Image TranslationTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Seed-to-Seed: Image Translation in Diffusion Seed Space
We introduce Seed-to-Seed Translation (StS), a novel approach for Image-to-Image Translation using diffusion models (DMs), aimed at translations that require close adherence to the structure of the source image. In contr…
Image-to-Image TranslationSTSTranslationNorm-guided latent space exploration for text-to-image generation
Text-to-image diffusion models show great potential in synthesizing a large variety of concepts in new compositions and scenarios. However, the latent space of initial seeds is still not well understood and its structure…
Image GenerationLong-tail LearningText to Image GenerationText-to-Image GenerationMemories of Forgotten Concepts
Diffusion models dominate the space of text-to-image generation, yet they may produce undesirable outputs, including explicit content or private data. To mitigate this, concept ablation techniques have been explored to l…
Image GenerationText to Image GenerationText-to-Image GenerationUnpaired Image-to-Image Translation via Latent Energy Transport
Image-to-image translation aims to preserve source contents while translating to discriminative target styles between two visual domains. Most works apply adversarial learning in the ambient image space, which could be c…
Image ReconstructionImage-to-Image TranslationTranslationSmoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image Translation
Image-to-Image (I2I) multi-domain translation models are usually evaluated also using the quality of their semantic interpolation results. However, state-of-the-art models frequently show abrupt changes in the image appe…
Image-to-Image TranslationTranslationUnsupervised Image-To-Image Translation