paper-with-me

Papers

Structure-Level Disentangled Diffusion for Few-Shot Chinese Font Generation

2026-02-21 · Jie Li, Suorong Yang, Jian Zhao, Furao Shen arxiv

Few-shot Chinese font generation aims to synthesize new characters in a target style using only a handful of reference images. Achieving accurate content rendering and faithful style transfer requires effective disentanglement between content and style. However, existing approaches achieve only feature-level disentanglement, allowing the generator to re-entangle these features, leading to content distortion and degraded style fidelity. We propose the Structure-Level Disentangled Diffusion Model (SLD-Font), which receives content and style information from two separate channels. SimSun-style images are used as content templates and concatenated with noisy latent features as the input. Style features extracted by a CLIP model from target-style images are integrated via cross-attention. Additionally, we train a Background Noise Removal module in the pixel space to remove background noise in complex stroke regions. Based on theoretical validation of disentanglement effectiveness, we introduce a parameter-efficient fine-tuning strategy that updates only the style-related modules. This allows the model to better adapt to new styles while avoiding overfitting to the reference images' content. We further introduce the Grey and OCR metrics to evaluate the content quality of generated characters. Experimental results show that SLD-Font achieves significantly higher style fidelity while maintaining comparable content accuracy to existing state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2602.18874

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningStyle Transfer

Similar Papers 제목 키워드 기반

The Curious Case of End Token: A Zero-Shot Disentangled Image Editing using CLIP

2024-06-01 · Hidir Yesiltepe, Yusuf Dalva, Pinar Yanardag

Diffusion models have become prominent in creating high-quality images. However, unlike GAN models celebrated for their ability to edit images in a disentangled manner, diffusion-based text-to-image models struggle to ac…

AttributeVideo Editing

RD-GAN: Few/Zero-Shot Chinese Character Style Transfer via Radical Decomposition and Rendering

2020-08-01 · ECCV 2020 8 · Yaoxiong Huang, Mengchao He, Lianwen Jin, Yongpan Wang

Style transfer has attracted much interest owing to its various applications. Compared with English character or general artistic style transfer, Chinese character style transfer remains a challenge owing to the large si…

Style TransferZero-shot Generalization

HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models

2025-05-10 · Shuhan Zhuang, Mengqi Huang, Fengyi Fu, Nan Chen 외

Visual text rendering, which aims to accurately integrate specified textual content within generated images, is critical for various applications such as commercial design. Despite recent advances, current methods strugg…

Text Generation

DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing

2023-12-11 · Haozhe Jia, Yan Li, Hengfei Cui, Di Xu 외

In this work, we focus on exploring explicit fine-grained control of generative facial image editing, all while generating faithful facial appearances and consistent semantic details, which however, is quite challenging …

DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model

2025-03-24 · Kangwei Liu, Junwu Liu, Yun Cao, Jinlin Guo 외

Recent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine…

DisentanglementFace GenerationTalking Face Generation