paper-with-me

Papers

SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder

2025-10-06 · Ronen Kamenetsky, Sara Dorfman, Daniel Garibi, Roni Paiss, Or Patashnik, Daniel Cohen-Or arxiv

Large-scale text-to-image diffusion models have become the backbone of modern image editing, yet text prompts alone do not offer adequate control over the editing process. Two properties are especially desirable: disentanglement, where changing one attribute does not unintentionally alter others, and continuous control, where the strength of an edit can be smoothly adjusted. We introduce a method for disentangled and continuous editing through token-level manipulation of text embeddings. The edits are applied by manipulating the embeddings along carefully chosen directions, which control the strength of the target attribute. To identify such directions, we employ a Sparse Autoencoder (SAE), whose sparse latent space exposes semantically isolated dimensions. Our method operates directly on text embeddings without modifying the diffusion process, making it model agnostic and broadly applicable to various image synthesis backbones. Experiments show that it enables intuitive and efficient manipulations with continuous control across diverse attributes and domains.

📄 PDF Abstract BibTeX arXiv:2510.05081

Code (0)

등록된 구현이 없습니다.

Tasks

Continuous ControlImage Editing

Similar Papers 제목 키워드 기반

All-in-One Slider for Attribute Manipulation in Diffusion Models

2025-08-26 · Weixin Ye, Hongguang Zhu, Wei Wang, Yahui Liu 외 arxiv

Text-to-image (T2I) diffusion models have made significant strides in generating high-quality images. However, progressively manipulating certain attributes of generated images to meet the desired user expectations remai…

Continuous Control

Token-to-Token Alignment of Text Embeddings for Semantic Blending

2026-06-22 · Saar Huberman, Ron Mokady, Or Patashnik, Daniel Cohen-Or arxiv

In modern generative models, images are specified and controlled through text prompts. In practice, images are generated from sequences of tokens derived from these prompts. However, the space of token sequences lacks a …

Semantic correspondenceSemantic SimilarityContinuous Control

Recolour What Matters: Region-Aware Colour Editing via Token-Level Diffusion

2026-03-19 · Yuqi Yang, Dongliang Chang, Yijia Ling, Ruoyi Du 외 arxiv

Colour is one of the most perceptually salient yet least controllable attributes in image generation. Although recent diffusion models can modify object colours from user instructions, their results often deviate from th…

Image Generation

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation

2026-06-03 · Dewei Zhou, Xinyu Huang, Xun Wang, Ji Xie 외 arxiv

Generative visual models fundamentally struggle with precise spatial control. This arises from a core disconnect: models can process textual descriptions of space but cannot directly map numerical coordinates onto the 2D…

Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models

2023-09-14 · James Burgess, Kuan-Chieh Wang, Serena Yeung-Levy

Text-to-image diffusion models generate impressive and realistic images, but do they learn to represent the 3D world from only 2D supervision? We demonstrate that yes, certain 3D scene representations are encoded in the …

Image GenerationNovel View SynthesisText to Image GenerationText-to-Image Generation