paper-with-me

Papers

Illiterate DALL$\cdot$E Learns to Compose

2021-09-29 · ICLR 2022 4 · Gautam Singh, Fei Deng, Sungjin Ahn

DALL$\cdot$E has shown an impressive ability of composition-based systematic generalization in image generation. This is possible because it utilizes the dataset of text-image pairs where the text provides the source of compositionality. Following this result, an important extending question is whether this compositionality can still be achieved even without conditioning on text. In this paper, we propose an architecture called $\textit{Slot2Seq}$ that achieves this text-free DALL$\cdot$E by learning compositional slot-based representations purely from images, an ability lacking in DALL$\cdot$E. Unlike existing object-centric representation models that decode pixels independently for each slot and each pixel location and compose them via mixture-based alpha composition, we propose to use the Image GPT decoder conditioned on the slots for a more flexible generation by capturing complex interaction among the pixels and the slots. In experiments, we show that this simple architecture achieves zero-shot generation of novel images without text and better quality in generation than the models based on mixture decoders.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage GenerationSystematic Generalization

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Weight Decay 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

Illiterate DALL-E Learns to Compose

2021-10-17 · Gautam Singh, Fei Deng, Sungjin Ahn

Although DALL-E has shown an impressive ability of composition-based systematic generalization in image generation, it requires the dataset of text-image pairs and the compositionality is provided by the text. In contras…

DecoderImage GenerationObjectSystematic Generalization

Improving dermatology classifiers across populations using images generated by large diffusion models

2022-11-23 · Luke W. Sagers, James A. Diao, Matthew Groh, Pranav Rajpurkar 외

Dermatological classification algorithms developed without sufficiently diverse training data may generalize poorly across populations. While intentional data collection and annotation offer the best means for improving …

An Annotated Reading of 'The Singer of Tales' in the LLM Era

2025-02-07 · Kush R. Varshney

The Parry-Lord oral-formulaic theory was a breakthrough in understanding how oral narrative poetry is learned, composed, and transmitted by illiterate bards. In this paper, we provide an annotated reading of the mechanis…

SneakyPrompt: Jailbreaking Text-to-image Generative Models

2023-05-20 · Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong 외

Text-to-image generative models such as Stable Diffusion and DALL$\cdot$E raise many ethical concerns due to the generation of harmful images such as Not-Safe-for-Work (NSFW) ones. To address these ethical concerns, safe…

Reinforcement Learning (RL)Semantic SimilaritySemantic Textual Similarity

Multiview Compressive Coding for 3D Reconstruction

2023-01-19 · CVPR 2023 1 · Chao-yuan Wu, Justin Johnson, Jitendra Malik, Christoph Feichtenhofer 외

A central goal of visual recognition is to understand objects and scenes from a single image. 2D recognition has witnessed tremendous progress thanks to large-scale learning and general-purpose representations. Comparati…

3D ReconstructionDecoderSelf-Supervised LearningSingle-View 3D Reconstruction