paper-with-me

홈 › Papers

TextWand: A Unified Framework for Scene Text Editing

2026-06-04 · Shuyu Wang, Zhile Guan, Hongxiu Chen, Yule Duan, Weiqi Li, Xin Shan, Ronggang Wang, Jian Zhang arxiv

We propose TextWand, a general-purpose framework that unifies scene text removal, generation, and replacement into a single model. By decomposing complex editing tasks into the atomic primitives of rendering and erasure, TextWand achieves precise control over both text appearance and background integrity. Specifically, we introduce a novel design, Overlay-Reference Positional Encoding (ORPE), to enforce pixel-level layout fidelity and exemplar-driven style control, alongside a new strategy, Region-Adaptive Suppression (RAS), to ensure clean text erasure. To address the absence of a comprehensive benchmark for general-purpose scene text editing among existing single-task datasets, we construct TextWand-Bench. Extensive experiments demonstrate that TextWand outperforms existing leading open-source and closed-source models by delivering superior text content accuracy, layout and style consistency, and overall image quality across scene text removal, generation and replacement tasks.

📄 PDF Abstract BibTeX arXiv:2606.05730

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unified Editing of Panorama, 3D Scenes, and Videos Through Disentangled Self-Attention Injection

2024-05-27 · Gihyun Kwon, Jangho Park, Jong Chul Ye

While text-to-image models have achieved impressive capabilities in image generation and editing, their application across various modalities often necessitates training separate models. Inspired by existing method of si…

Image GenerationVideo Editing

JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space

2026-06-11 · Xinnan Zhu, Ruijie Xu, Jiayu Ying, Daoguo Dong 외 arxiv

Existing 3D scene editing methods typically rely on per-scene optimization over explicit 3D representations or cascaded edit-and-reconstruct pipelines, resulting in high test-time cost, limited 3D awareness, and structur…

3D scene Editing

3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting

2024-05-28 · Qihang Zhang, Yinghao Xu, Chaoyang Wang, Hsin-Ying Lee 외

Scene image editing is crucial for entertainment, photography, and advertising design. Existing methods solely focus on either 2D individual object or 3D global scene editing. This results in a lack of a unified approach…

3D geometryDisentanglement

Native3D: End-to-End 3D Scene Generation via Unified Mesh-Texture Modeling and Semantic Alignment

2026-06-05 · Yibo Liu, Ziwei Zhang, Haozhou Pang, Menghao Li 외 arxiv

This paper presents Native3D, the first end-to-end 3D scene generation framework that completely bypasses 2D intermediate representations. Traditional approaches typically require adapting 3D representations to the 2D do…

Contrastive LearningDomain AdaptationScene Generation3D scene Editing

SimGraph: A Unified Framework for Scene Graph-Based Image Generation and Editing

2026-01-29 · Thanh-Nhan Vo, Trong-Thuan Nguyen, Tam V. Nguyen, Minh-Triet Tran arxiv

Recent advancements in Generative Artificial Intelligence (GenAI) have significantly enhanced the capabilities of both image generation and editing. However, current approaches often treat these tasks separately, leading…

Image Generation