paper-with-me

홈 › Papers

Generating Annotated High-Fidelity Images Containing Multiple Coherent Objects

2020-06-22 · Bryan G. Cardenas, Devanshu Arya, Deepak K. Gupta

Recent developments related to generative models have made it possible to generate diverse high-fidelity images. In particular, layout-to-image generation models have gained significant attention due to their capability to generate realistic complex images containing distinct objects. These models are generally conditioned on either semantic layouts or textual descriptions. However, unlike natural images, providing auxiliary information can be extremely hard in domains such as biomedical imaging and remote sensing. In this work, we propose a multi-object generation framework that can synthesize images with multiple objects without explicitly requiring their contextual information during the generation process. Based on a vector-quantized variational autoencoder (VQ-VAE) backbone, our model learns to preserve spatial coherency within an image as well as semantic coherency between the objects and the background through two powerful autoregressive priors: PixelSNAIL and LayoutPixelSNAIL. While the PixelSNAIL learns the distribution of the latent encodings of the VQ-VAE, the LayoutPixelSNAIL is used to specifically learn the semantic distribution of the objects. An implicit advantage of our approach is that the generated samples are accompanied by object-level annotations. We demonstrate how coherency and fidelity are preserved with our method through experiments on the Multi-MNIST and CLEVR datasets; thereby outperforming state-of-the-art multi-object generative methods. The efficacy of our approach is demonstrated through application on medical imaging datasets, where we show that augmenting the training set with generated samples using our approach improves the performance of existing models.

📄 PDF Abstract BibTeX arXiv:2006.12150

Code (1)

Cynetics/MSGNet 공식 구현 pytorch

Tasks

Image GenerationLayout-to-Image GenerationVocal Bursts Intensity Prediction

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…
Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Cellcounter: a deep learning framework for high-fidelity spatial localization of neurons

2021-03-18 · Tamal Batabyal, Aijaz Ahmad Naik, Daniel Weller, Jaideep Kapur

Many neuroscientific applications require robust and accurate localization of neurons. It is still an unsolved problem because of the enormous variation in intensity, texture, spatial overlap, morphology and background a…

Self-Learning

NURBGen: High-Fidelity Text-to-CAD Generation through LLM-Driven NURBS Modeling

2025-11-09 · Muhammad Usama, Mohammad Sadil Khan, Didier Stricker, Muhammad Zeshan Afzal arxiv

Generating editable 3D CAD models from natural language remains challenging, as existing text-to-CAD systems either produce meshes or rely on scarce design-history data. We present NURBGen, the first framework to generat…

Re-Imagen: Retrieval-Augmented Text-to-Image Generator

2022-09-29 · Wenhu Chen, Hexiang Hu, Chitwan Saharia, William W. Cohen

Research on text-to-image generation has witnessed significant progress in generating diverse and photo-realistic images, driven by diffusion and auto-regressive models trained on large-scale image-text data. Though stat…

Image GenerationImage-text RetrievalRetrievalText Retrieval+2

ObjectComposer: Consistent Generation of Multiple Objects Without Fine-tuning

2023-10-10 · Alec Helbling, Evan Montoya, Duen Horng Chau

Recent text-to-image generative models can generate high-fidelity images from text prompts. However, these models struggle to consistently generate the same objects in different contexts with the same appearance. Consist…

Semantic and Visual Crop-Guided Diffusion Models for Heterogeneous Tissue Synthesis in Histopathology

2025-09-22 · Saghir Alfasly, Wataru Uegami, MD Enamul Hoq, Ghazal Alabtah 외 arxiv

Synthetic data generation in histopathology faces unique challenges: preserving tissue heterogeneity, capturing subtle morphological features, and scaling to unannotated datasets. We present a latent diffusion model that…

Synthetic Data GenerationSemantic Segmentation