Modulated Contrast for Versatile Image Synthesis
Perceiving the similarity between images has been a long-standing and fundamental problem underlying various visual generation tasks. Predominant approaches measure the inter-image distance by computing pointwise absolute deviations, which tends to estimate the median of instance distributions and leads to blurs and artifacts in the generated images. This paper presents MoNCE, a versatile metric that introduces image contrast to learn a calibrated metric for the perception of multifaceted inter-image distances. Unlike vanilla contrast which indiscriminately pushes negative samples from the anchor regardless of their similarity, we propose to re-weight the pushing force of negative samples adaptively according to their similarity to the anchor, which facilitates the contrastive learning from informative negative samples. Since multiple patch-level contrastive objectives are involved in image distance measurement, we introduce optimal transport in MoNCE to modulate the pushing force of negative samples collaboratively across multiple contrastive objectives. Extensive experiments over multiple image translation tasks show that the proposed MoNCE outperforms various prevailing metrics substantially. The code is available at https://github.com/fnzhan/MoNCE.
Code (1)
Tasks
Contrastive LearningImage GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Universal Silicon Microwave Photonic Spectral Shaper
Optical modulation plays arguably the utmost important role in microwave photonic (MWP) systems. Precise synthesis of modulated optical spectra dictates virtually all aspects of MWP system quality including loss, noise f…
BlockingSpiS-GAN: Spiral-Modulated Handwriting Synthesis with Star Operation
Training robust handwriting recognition (HTR) systems requires massive amounts of annotated data, which is often difficult to acquire. While synthetic handwriting generation offers a practical solution to expand training…
Handwriting RecognitionGaussian Process Modulated Cox Processes under Linear Inequality Constraints
Gaussian process (GP) modulated Cox processes are widely used to model point patterns. Existing approaches require a mapping (link function) between the unconstrained GP and the positive intensity function. This commonly…
Point ProcessesVersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to …
AI AgentImage GenerationTop-Down Visual Attention from Analysis by Synthesis
Current attention algorithms (e.g., self-attention) are stimulus-driven and highlight all the salient objects in an image. However, intelligent agents like humans often guide their attention based on the high-level task …
RetrievalSemantic SegmentationVisual Question Answering (VQA)