paper-with-me

Papers

Cascading Modular Network (CAM-Net) for Multimodal Image Synthesis

2021-06-16 · Shichong Peng, Alireza Moazeni, Ke Li

Deep generative models such as GANs have driven impressive advances in conditional image synthesis in recent years. A persistent challenge has been to generate diverse versions of output images from the same input image, due to the problem of mode collapse: because only one ground truth output image is given per input image, only one mode of the conditional distribution is modelled. In this paper, we focus on this problem of multimodal conditional image synthesis and build on the recently proposed technique of Implicit Maximum Likelihood Estimation (IMLE). Prior IMLE-based methods required different architectures for different tasks, which limit their applicability, and were lacking in fine details in the generated images. We propose CAM-Net, a unified architecture that can be applied to a broad range of tasks. Additionally, it is capable of generating convincing high frequency details, achieving a reduction of the Frechet Inception Distance (FID) by up to 45.3% compared to the baseline.

📄 PDF Abstract BibTeX arXiv:2106.09015

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Modular Control of Discrete Event System for Modeling and Mitigating Power System Cascading Failures

2025-04-10 · Wasseem Al-Rousan, Caisheng Wang, Feng Lin

Cascading failures in power systems caused by sequential tripping of components are a serious concern as they can lead to complete or partial shutdowns, disrupting vital services and causing damage and inconvenience. In …

CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning

2024-10-15 · Qingqing Cao, Mahyar Najibi, Sachin Mehta

Pretraining robust vision or multimodal foundation models (e.g., CLIP) relies on large-scale datasets that may be noisy, potentially misaligned, and have long-tail distributions. Previous works have shown promising resul…

Image-text RetrievalText Retrievalzero-shot-classificationZero-Shot Learning

ChartSync: A Benchmark for Visuo-Logical Cascading Chart Editing

2026-07-11 · Jiakang Yu, Yixuan Chai, Tianci Wang, Rihui Jin 외 arxiv

Generative image editing models struggle with structured statistical charts when data modifications require geometric synchronization. We formalize this task as Visuo-Logical Cascading Editing (VLCE). However, existing m…

Image Editing

On the Perils of Cascading Robust Classifiers

2022-06-01 · Ravi Mangal, Zifan Wang, Chi Zhang, Klas Leino 외

Ensembling certifiably robust neural networks is a promising approach for improving the \emph{certified robust accuracy} of neural models. Black-box ensembles that assume only query-access to the constituent models (and …

Adversarial Attack

LetsTalk: Latent Diffusion Transformer for Talking Video Synthesis

2024-11-24 · Haojie Zhang, Zhihao Liang, Ruibo Fu, Zhengqi Wen 외

Portrait image animation using audio has rapidly advanced, enabling the creation of increasingly realistic and expressive animated faces. The challenges of this multimodality-guided video generation task involve fusing v…

DiversityImage AnimationVideo Generation