paper-with-me

Papers

InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization

2025-11-18 · Daniel Gilo, Or Litany arxiv

We address the task of multi-view image editing from sparse input views, where the inputs can be seen as a mix of images capturing the scene from different viewpoints. The goal is to modify the scene according to a textual instruction while preserving consistency across all views. Existing methods, based on per-scene neural fields or temporal attention mechanisms, struggle in this setting, often producing artifacts and incoherent edits. We propose InstructMix2Mix (I-Mix2Mix), a framework that distills the editing capabilities of a 2D diffusion model into a pretrained multi-view diffusion model, leveraging its data-driven 3D prior for cross-view consistency. A key contribution is replacing the conventional neural field consolidator in Score Distillation Sampling (SDS) with a multi-view diffusion student, which requires novel adaptations: incremental student updates across timesteps, a specialized teacher noise scheduler to prevent degeneration, and an attention modification that enhances cross-view coherence without additional cost. Experiments demonstrate that I-Mix2Mix significantly improves multi-view consistency while maintaining high per-frame edit quality.

📄 PDF Abstract BibTeX arXiv:2511.14899

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing

2025-05-31 · Yang Zheng, Mengqi Huang, Nan Chen, Zhendong Mao

Text-guided 3D editing aims to precisely edit semantically relevant local 3D regions, which has significant potential for various practical applications ranging from 3D games to film production. Existing methods typicall…

FluSplat: Sparse-View 3D Editing without Test-Time Optimization

2026-04-21 · Haitao Huang, Shin-Fang Chng, Huangying Zhan, Qingan Yan 외 arxiv

Recent advances in text-guided image editing and 3D Gaussian Splatting (3DGS) have enabled high-quality 3D scene manipulation. However, existing pipelines rely on iterative edit-and-fit optimization at test time, alterna…

3D Reconstruction3D scene EditingImage Editing

Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization

2025-08-20 · Canyu Zhao, Xiaoman Li, Tianjian Feng, Zhiyue Zhao 외 arxiv

We introduce Tinker, a versatile framework for high-fidelity 3D editing that operates in both one-shot and few-shot regimes without any per-scene finetuning. Unlike prior techniques that demand extensive per-scene optimi…

Coupled Diffusion Sampling for Training-Free Multi-View Image Editing

2025-10-16 · Hadi Alzayer, Yunzhi Zhang, Chen Geng, Jia-Bin Huang 외 arxiv

We present an inference-time diffusion sampling method to perform multi-view consistent image editing using pre-trained 2D image editing models. These models can independently produce high-quality edits for each image in…

Image Editing

3D-Consistent Multi-View Editing by Correspondence Guidance

2025-11-27 · Josef Bengtson, David Nilsson, Dong In Lee, Yaroslava Lochman 외 arxiv

Recent advancements in diffusion and flow models have greatly improved text-based image editing, yet methods that edit images independently often produce geometrically and photometrically inconsistent results across diff…

Text-based Image Editing