paper-with-me

홈 › Papers

DreamOmni3: Scribble-based Editing and Generation

2025-12-27 · Bin Xia, Bohao Peng, Jiyang Liu, Sitong Wu, Jingyao Li, Junjia Huang, Xu Zhao, Yitong Wang, Ruihang Chu, Bei Yu, Jiaya Jia arxiv

Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text prompts for instruction-based editing and generation, but language often fails to capture users intended edit locations and fine-grained visual details. To this end, we propose two tasks: scribble-based editing and generation, that enables more flexible creation on graphical user interface (GUI) combining user textual, images, and freehand sketches. We introduce DreamOmni3, tackling two challenges: data creation and framework design. Our data synthesis pipeline includes two parts: scribble-based editing and generation. For scribble-based editing, we define four tasks: scribble and instruction-based editing, scribble and multimodal instruction-based editing, image fusion, and doodle editing. Based on DreamOmni2 dataset, we extract editable regions and overlay hand-drawn boxes, circles, doodles or cropped image to construct training data. For scribble-based generation, we define three tasks: scribble and instruction-based generation, scribble and multimodal instruction-based generation, and doodle generation, following similar data creation pipelines. For the framework, instead of using binary masks, which struggle with complex edits involving multiple scribbles, images, and instructions, we propose a joint input scheme that feeds both the original and scribbled source images into the model, using different colors to distinguish regions and simplify processing. By applying the same index and position encodings to both images, the model can precisely localize scribbled regions while maintaining accurate editing. Finally, we establish comprehensive benchmarks for these tasks to promote further research. Experimental results demonstrate that DreamOmni3 achieves outstanding performance, and models and code will be publicly released.

📄 PDF Abstract BibTeX arXiv:2512.22525

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DreamOmni: Unified Image Generation and Editing

2024-12-22 · CVPR 2025 1 · Bin Xia, Yuechen Zhang, Jingyao Li, Chengyao Wang 외

Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster synergistic benefits across different t…

Image Generation

DreamOmni2: Multimodal Instruction-based Editing and Generation

2025-10-08 · Bin Xia, Bohao Peng, Yuechen Zhang, Junjia Huang 외 arxiv

Recent advancements in instruction-based image editing and subject-driven generation have garnered significant attention, yet both tasks still face limitations in meeting practical user needs. Instruction-based editing r…

Image Editing

ScribbleSense: Generative Scribble-Based Texture Editing with Intent Prediction

2026-01-30 · Yudi Zhang, Yeming Geng, Lei Zhang arxiv

Interactive 3D model texture editing presents enhanced opportunities for creating 3D assets, with freehand drawing style offering the most intuitive experience. However, existing methods primarily support sketch-based in…

Image Generation

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking

2026-05-25 · Mingyi Xu, Jinpeng Lin, Min Zhou, Tiezheng Ge 외 arxiv

Scribble-guided image editing allows users to combine simple scribble annotations with text prompts to specify both where and how an image should be edited, enabling flexible interaction with precise spatial control. How…

Domain GeneralizationImage Editing

ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text

2026-05-01 · Anya Ji, George Ma, Téa Wright, Yiming Zhang 외 arxiv

Recent progress in generative models has significantly advanced image editing capabilities, yet precise and intuitive user control remains difficult. Specifically, users often struggle to communicate both exact spatial l…

Image Editing