paper-with-me

Papers

Generative Adversarial Image Synthesis with Decision Tree Latent Controller

2018-05-27 · CVPR 2018 6 · Takuhiro Kaneko, Kaoru Hiramatsu, Kunio Kashino

This paper proposes the decision tree latent controller generative adversarial network (DTLC-GAN), an extension of a GAN that can learn hierarchically interpretable representations without relying on detailed supervision. To impose a hierarchical inclusion structure on latent variables, we incorporate a new architecture called the DTLC into the generator input. The DTLC has a multiple-layer tree structure in which the ON or OFF of the child node codes is controlled by the parent node codes. By using this architecture hierarchically, we can obtain the latent space in which the lower layer codes are selectively used depending on the higher layer ones. To make the latent codes capture salient semantic features of images in a hierarchically disentangled manner in the DTLC, we also propose a hierarchical conditional mutual information regularization and optimize it with a newly defined curriculum learning method that we propose as well. This makes it possible to discover hierarchically interpretable representations in a layer-by-layer manner on the basis of information gain by only using a single DTLC-GAN model. We evaluated the DTLC-GAN on various datasets, i.e., MNIST, CIFAR-10, Tiny ImageNet, 3D Faces, and CelebA, and confirmed that the DTLC-GAN can learn hierarchically interpretable representations with either unsupervised or weakly supervised settings. Furthermore, we applied the DTLC-GAN to image-retrieval tasks and showed its effectiveness in representation learning.

📄 PDF Abstract BibTeX arXiv:1805.10603

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Adversarial NetworkImage GenerationImage RetrievalRepresentation LearningRetrieval

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks

2017-10-19 · Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang 외

Although Generative Adversarial Networks (GANs) have shown remarkable success in various tasks, they still face challenges in generating high quality images. In this paper, we propose Stacked Generative Adversarial Netwo…

Generative Adversarial NetworkImage GenerationText-to-Image Generation

From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis

2025-09-29 · Khawlah Bajbaa, Abbas Anwar, Muhammad Saqib, Hafeez Anwar 외 arxiv

Street view imagery has become an essential source for geospatial data collection and urban analytics, enabling the extraction of valuable insights that support informed decision-making. However, synthesizing street-view…

Example-Guided Style-Consistent Image Synthesis From Semantic Labeling

2019-06-01 · CVPR 2019 6 · Miao Wang, Guo-Ye Yang, Ruilong Li, Run-Ze Liang 외

Example-guided image synthesis aims to synthesize an image from a semantic label map and an exemplary image indicating style. We use the term "style" in this problem to refer to implicit characteristics of images, for e…

Image GenerationScene Segmentation

Re-designing cities with conditional adversarial networks

2021-04-08 · Mohamed R. Ibrahim, James Haworth, Nicola Christie

This paper introduces a conditional generative adversarial network to redesign a street-level image of urban scenes by generating 1) an urban intervention policy, 2) an attention map that localises where intervention is …

Generative Adversarial NetworkGPUImage GenerationImage-to-Image Translation+2

Video-to-Video Synthesis

2018-08-20 · NeurIPS 2018 12 · Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu 외

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely de…

2kSemantic SegmentationVideo derainingVideo Prediction+1