paper-with-me

홈 › Papers

DreamOmni2: Multimodal Instruction-based Editing and Generation

2025-10-08 · Bin Xia, Bohao Peng, Yuechen Zhang, Junjia Huang, Jiyang Liu, Jingyao Li, Haoru Tan, Sitong Wu, Chengyao Wang, Yitong Wang, Xinglong Wu, Bei Yu, Jiaya Jia arxiv

Recent advancements in instruction-based image editing and subject-driven generation have garnered significant attention, yet both tasks still face limitations in meeting practical user needs. Instruction-based editing relies solely on language instructions, which often fail to capture specific editing details, making reference images necessary. Meanwhile, subject-driven generation is limited to combining concrete objects or people, overlooking broader, abstract concepts. To address these challenges, we propose two novel tasks: multimodal instruction-based editing and generation. These tasks support both text and image instructions and extend the scope to include both concrete and abstract concepts, greatly enhancing their practical applications. We introduce DreamOmni2, tackling two primary challenges: data creation and model framework design. Our data synthesis pipeline consists of three steps: (1) using a feature mixing method to create extraction data for both abstract and concrete concepts, (2) generating multimodal instruction-based editing training data using the editing and extraction models, and (3) further applying the extraction model to create training data for multimodal instruction-based editing. For the framework, to handle multi-image input, we propose an index encoding and position encoding shift scheme, which helps the model distinguish images and avoid pixel confusion. Additionally, we introduce joint training with the VLM and our generation/editing model to better process complex instructions. In addition, we have proposed comprehensive benchmarks for these two new tasks to drive their development. Experiments show that DreamOmni2 has achieved impressive results. Models and codes will be released.

📄 PDF Abstract BibTeX arXiv:2510.06679

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

DreamOmni3: Scribble-based Editing and Generation

2025-12-27 · Bin Xia, Bohao Peng, Jiyang Liu, Sitong Wu 외 arxiv

Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text prompts for instruction-based editing and generation, but language often…

DreamOmni: Unified Image Generation and Editing

2024-12-22 · CVPR 2025 1 · Bin Xia, Yuechen Zhang, Jingyao Li, Chengyao Wang 외

Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster synergistic benefits across different t…

Image Generation

MIGE: A Unified Framework for Multimodal Instruction-Based Image Generation and Editing

2025-02-28 · Xueyun Tian, Wei Li, Bingbing Xu, Yige Yuan 외

Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high…

Image GenerationTransfer Learning

UniVideo: Unified Understanding, Generation, and Editing for Videos

2025-10-09 · Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang 외 arxiv

Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends un…

Video GenerationText GenerationStyle TransferImage Editing

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

2026-02-10 · Jialun Liu, Tian Li, Xiao Cao, Yukuo Ma 외 arxiv

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, l…

Text-to-Video Generation