paper-with-me

홈 › Papers

GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation

2023-03-31 · Jian Ma, Mingjun Zhao, Chen Chen, Ruichen Wang, Di Niu, Haonan Lu, Xiaodong Lin

Recent breakthroughs in the field of language-guided image generation have yielded impressive achievements, enabling the creation of high-quality and diverse images based on user instructions.Although the synthesis performance is fascinating, one significant limitation of current image generation models is their insufficient ability to generate text coherently within images, particularly for complex glyph structures like Chinese characters. To address this problem, we introduce GlyphDraw, a general learning framework aiming to endow image generation models with the capacity to generate images coherently embedded with text for any specific language.We first sophisticatedly design the image-text dataset's construction strategy, then build our model specifically on a diffusion-based image generator and carefully modify the network structure to allow the model to learn drawing language characters with the help of glyph and position information.Furthermore, we maintain the model's open-domain image synthesis capability by preventing catastrophic forgetting by using parameter-efficient fine-tuning techniques.Extensive qualitative and quantitative experiments demonstrate that our method not only produces accurate language characters as in prompts, but also seamlessly blends the generated text into the background.Please refer to our \href{https://1073521013.github.io/glyph-draw.github.io/}{project page}. \end{abstract}

📄 PDF Abstract BibTeX arXiv:2303.17870

Code (3)

OPPO-Mente-Lab/GlyphDraw 공식 구현 pytorch
OPPO-Mente-Lab/Subject-Diffusion pytorch
aigtext/glyphcontrol-release pytorch

Tasks

Image GenerationOptical Character Recognition (OCR)parameter-efficient fine-tuningText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language Models

2024-07-02 · Jian Ma, Yonglin Deng, Chen Chen, Haonan Lu 외

Posters play a crucial role in marketing and advertising by enhancing visual communication and brand visibility, making significant contributions to industrial design. With the latest advancements in controllable T2I dif…

Marketing

Voxel-Mesh Hybrid Representation for Real-Time View Synthesis

2024-03-11 · Chenhao Zhang, Yongyang Zhou, Lei Zhang

The neural radiance fields (NeRF) have emerged as a prominent methodology for synthesizing realistic images of novel views. While neural radiance representations based on voxels or mesh individually offer distinct advant…

NeRFNeural Rendering

High-Fidelity 3D Head Avatars Reconstruction through Spatially-Varying Expression Conditioned Neural Radiance Field

2023-10-10 · Minghan Qin, Yifan Liu, Yuelang Xu, Xiaochen Zhao 외

One crucial aspect of 3D head avatar reconstruction lies in the details of facial expressions. Although recent NeRF-based photo-realistic 3D head avatar methods achieve high-quality avatar rendering, they still encounter…

NeRF

TextureSplat: Per-Primitive Texture Mapping for Reflective Gaussian Splatting

2025-06-16 · Mae Younes, Adnane Boukhayma

Gaussian Splatting have demonstrated remarkable novel view synthesis performance at high rendering frame rates. Optimization-based inverse rendering within complex capture scenarios remains however a challenging problem.…

GPUInverse RenderingNovel View Synthesis

Lacunarity Pooling Layers for Plant Image Classification using Texture Analysis

2024-04-25 · Akshatha Mohan, Joshua Peeples

Pooling layers (e.g., max and average) may overlook important information encoded in the spatial arrangement of pixel intensity and/or feature values. We propose a novel lacunarity pooling layer that aims to capture the …

image-classificationImage ClassificationTexture Classification