paper-with-me

홈 › Papers

DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models

2026-07-14 · Francesco Taioli, Daniel Coelho, Iaroslav Melekhov, Roberto Alcover-Couso, Jose Miguel Grande Saiz, Virginia Fernandez Arguedas, Artur Bekasov arxiv

Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions. First, we introduce ABO-Edit, a dataset specifically designed to study object consistency, comprising over 12,000 triplets of source images, editing prompts, and high-quality target images rendered from artist-designed 3D assets, with multi-view coverage and human-verified quality control. Second, we uncover an overlooked property of image-editing rectified flow models: the conditioning embedding space, not directly supervised during training, encodes a prediction of the final generated image even at high noise levels. Third, exploiting this finding, we propose FlowMirror, a parameter-free auxiliary loss that supervises this conditioning embedding space. Without architectural changes, our method improves generation quality across several metrics over baselines.

📄 PDF Abstract BibTeX arXiv:2607.12539

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

OR-NeRF: Object Removing from 3D Scenes Guided by Multiview Segmentation with Neural Radiance Fields

2023-05-17 · Youtan Yin, Zhoujie Fu, Fan Yang, Guosheng Lin

The emergence of Neural Radiance Fields (NeRF) for novel view synthesis has increased interest in 3D scene editing. An essential task in editing is removing objects from a scene while ensuring visual reasonability and mu…

3D geometry3D scene EditingNeRFNovel View Synthesis+1

DiffusionAtlas: High-Fidelity Consistent Diffusion Video Editing

2023-12-05 · Shao-Yu Chang, Hwann-Tzong Chen, Tyng-Luh Liu

We present a diffusion-based video editing framework, namely DiffusionAtlas, which can achieve both frame consistency and high fidelity in editing video object appearance. Despite the success in image editing, diffusion …

ObjectVideo Editing

3DSwapping: Texture Swapping For 3D Object From Single Reference Image

2025-03-24 · Xiao Cao, Beibei Lin, Bo wang, Zhiyong Huang 외

3D texture swapping allows for the customization of 3D object textures, enabling efficient and versatile visual transformations in 3D editing. While no dedicated method exists, adapted 2D editing and text-driven 3D editi…

ELDiff: When Evidential Learning Meets Text-to-Image Diffusion

2026-06-18 · Qingtao Pan, Kai Ye, Zhihao Dou, Bing Ji 외 arxiv

In multi-object text-to-image (T2I) diffusion, ensuring semantic consistency between textual prompts and generated visual content is crucial for image synthesis. However, such consistency constraint is often underemphasi…

Object Segmentation

Text-to-Audio Generation Synchronized with Videos

2024-03-08 · Shentong Mo, Jing Shi, Yapeng Tian

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models…

AudioCapsAudio GenerationContrastive Learning