paper-with-me

홈 › Papers

InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following

2023-12-11 · Shufan Li, Harkanwar Singh, Aditya Grover

The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explored extending controllability in two directions: instruction tuning with text-based prompts and multi-modal conditioning. However, these works make one or more unnatural assumptions on the number and/or type of modality inputs used to express controllability. We propose InstructAny2Pix, a flexible multi-modal instruction-following system that enables users to edit an input image using instructions involving audio, images, and text. InstructAny2Pix consists of three building blocks that facilitate this capability: a multi-modal encoder that encodes different modalities such as images and audio into a unified latent space, a diffusion model that learns to decode representations in this latent space into images, and a multi-modal LLM that can understand instructions involving multiple images and audio pieces and generate a conditional embedding of the desired output, which can be used by the diffusion decoder. Additionally, to facilitate training efficiency and improve generation quality, we include an additional refinement prior module that enhances the visual quality of LLM outputs. These designs are critical to the performance of our system. We demonstrate that our system can perform a series of novel instruction-guided editing tasks. The code is available at https://github.com/jacklishufan/InstructAny2Pix.git

📄 PDF Abstract BibTeX arXiv:2312.06738

Code (1)

jacklishufan/instructany2pix 공식 구현 pytorch

Tasks

DecoderInstruction Following

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MSRAMIE: Multimodal Structured Reasoning Agent for Multi-instruction Image Editing

2026-03-17 · Zhaoyuan Qiu, Ken Chen, Xiangwei Wang, Yu Xia 외 arxiv

Existing instruction-based image editing models perform well with simple, single-step instructions but degrade in realistic scenarios that involve multiple, lengthy, and interdependent directives. A main cause is the sca…

Instruction FollowingMultimodal ReasoningImage Editing

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

2026-02-10 · Jialun Liu, Tian Li, Xiao Cao, Yukuo Ma 외 arxiv

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, l…

Text-to-Video Generation

MIGE: A Unified Framework for Multimodal Instruction-Based Image Generation and Editing

2025-02-28 · Xueyun Tian, Wei Li, Bingbing Xu, Yige Yuan 외

Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high…

Image GenerationTransfer Learning

DreamOmni3: Scribble-based Editing and Generation

2025-12-27 · Bin Xia, Bohao Peng, Jiyang Liu, Sitong Wu 외 arxiv

Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text prompts for instruction-based editing and generation, but language often…

VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning

2025-04-03 · Xianwei Zhuang, Yuxin Xie, Yufan Deng, Dongchao Yang 외

In this work, we present VARGPT-v1.1, an advanced unified visual autoregressive model that builds upon our previous framework VARGPT. The model preserves the dual paradigm of next-token prediction for visual understandin…

Image GenerationInstruction Following