paper-with-me

홈 › Papers

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

2024-12-10 · CVPR 2025 1 · Xi Chen, Zhifei Zhang, He Zhang, Yuqian Zhou, Soo Ye Kim, Qing Liu, Yijun Li, Jianming Zhang, Nanxuan Zhao, Yilin Wang, Hui Ding, Zhe Lin, Hengshuang Zhao

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation models that effectively balance consistency and variation across frames, we propose a unifying approach that treats image-level tasks as discontinuous video generation. Specifically, we treat varying numbers of input and output images as frames, enabling seamless support for tasks such as image generation, editing, customization, composition, etc. Although designed for image-level tasks, we leverage videos as a scalable source for universal supervision. UniReal learns world dynamics from large-scale videos, demonstrating advanced capability in handling shadows, reflections, pose variation, and object interaction, while also exhibiting emergent capability for novel applications.

📄 PDF Abstract BibTeX arXiv:2412.07774

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationVideo Generation

Similar Papers 제목 키워드 기반

Large-Scale Universal Defect Generation: Foundation Models and Datasets

2026-04-10 · Yuanting Fan, Jun Liu, Bin-Bin Gao, Xiaochen Chen 외 arxiv

Existing defect/anomaly generation methods often rely on few-shot learning, which overfits to specific defect categories due to the lack of large-scale paired defect editing data. This issue is aggravated by substantial …

Multi-class Anomaly DetectionFew-Shot Learning

DiffUTE: Universal Text Editing Diffusion Model

2023-05-18 · NeurIPS 2023 11 · Haoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan 외

Diffusion model based language-guided image editing has achieved great success recently. However, existing state-of-the-art diffusion models struggle with rendering correct text and text style during generation. To tackl…

modelSelf-Supervised Learning

OmniCreator: Self-Supervised Unified Generation with Universal Editing

2024-12-03 · Haodong Chen, Lan Wang, Harry Yang, Ser-Nam Lim

We introduce OmniCreator, a novel framework that can conduct text-prompted unified (image+video) generation as well as editing all in one place. OmniCreator acquires generative and universal editing capabilities in a sel…

DenoisingSemantic correspondenceVideo EditingVideo Generation

VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On

2026-03-12 · Xiaoye Liang, Zhiyuan Qu, Mingye Zou, Jiaxin Liu 외 arxiv

As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing specialized VTON models. Meanwhile, universal multi-reference image editing…

Virtual Try-onImage Editing

Eta Inversion: Designing an Optimal Eta Function for Diffusion-based Real Image Editing

2024-03-14 · Wonjun Kang, Kevin Galim, Hyung Il Koo

Diffusion models have achieved remarkable success in the domain of text-guided image generation and, more recently, in text-guided image editing. A commonly adopted strategy for editing real images involves inverting the…

Image Generationtext-guided-image-editing