paper-with-me

홈 › Papers

NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion

2021-11-24 · Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, Nan Duan

This paper presents a unified multimodal pre-trained model called N\"UWA that can generate new or manipulate existing visual data (i.e., images and videos) for various visual synthesis tasks. To cover language, image, and video at the same time for different scenarios, a 3D transformer encoder-decoder framework is designed, which can not only deal with videos as 3D data but also adapt to texts and images as 1D and 2D data, respectively. A 3D Nearby Attention (3DNA) mechanism is also proposed to consider the nature of the visual data and reduce the computational complexity. We evaluate N\"UWA on 8 downstream tasks. Compared to several strong baselines, N\"UWA achieves state-of-the-art results on text-to-image generation, text-to-video generation, video prediction, etc. Furthermore, it also shows surprisingly good zero-shot capabilities on text-guided image and video manipulation tasks. Project repo is https://github.com/microsoft/NUWA.

📄 PDF Abstract BibTeX arXiv:2111.12417

Code (1)

lucidrains/nuwa-pytorch pytorch

Tasks

DecoderImage GenerationText to Image GenerationText-to-Image GenerationText-to-Video GenerationVideo GenerationVideo Prediction

Methods 이 논문이 사용한 방법론

1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
1cycle 설명 없음
LAMB LAMB is a a layerwise adaptive large batch optimization technique. It provides a strategy for adapting the learning rate in large batch settings. LAMB uses…
(USA Guide) 설명 없음
Adam 설명 없음
1-bit Adam 1-bit Adam is a stochastic optimization technique that is a variant of…

Similar Papers 제목 키워드 기반

Generative Adversarial Networks for Image and Video Synthesis: Algorithms and Applications

2020-08-06 · Ming-Yu Liu, Xun Huang, Jiahui Yu, Ting-Chun Wang 외

The generative adversarial network (GAN) framework has emerged as a powerful tool for various image and video synthesis tasks, allowing the synthesis of visual content in an unconditional or input-conditional manner. It …

Generative Adversarial NetworkNeural RenderingTranslation

BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation

2026-03-26 · Yan Li, Zezi Zeng, Ziwei Zhou, Xin Gao 외 arxiv

Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly focus on natural image synthesis and fail…

Image Generation

PARASOL: Parametric Style Control for Diffusion Image Synthesis

2023-03-11 · Gemma Canet Tarrés, Dan Ruta, Tu Bui, John Collomosse

We propose PARASOL, a multi-modal synthesis model that enables disentangled, parametric control of the visual style of the image by jointly conditioning synthesis on both content and a fine-grained visual style embedding…

Image Generation

Effective Training Data Synthesis for Improving MLLM Chart Understanding

2025-08-08 · Yuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li 외 arxiv

Being able to effectively read scientific plots, or chart understanding, is a central part toward building effective agents for science. However, existing multimodal large language models (MLLMs), especially open-source …

Exemplar-based Pattern Synthesis with Implicit Periodic Field Network

2022-04-04 · CVPR 2022 1 · Haiwei Chen, Jiayi Liu, Weikai Chen, Shichen Liu 외

Synthesis of ergodic, stationary visual patterns is widely applicable in texturing, shape modeling, and digital content creation. The wide applicability of this technique thus requires the pattern synthesis approaches to…

Generative Adversarial NetworkTexture Synthesis