paper-with-me

홈 › Papers

SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design

2025-11-17 · Yunjie Yu, Jingchen Wu, Junchen Zhu, Chunze Lin, Guibin Chen arxiv

Artistic design, particularly poster design, often demands rapid yet precise modification of textual content while preserving visual harmony and typographic intent, especially across diverse font styles. Although modern image editing models have grown increasingly powerful, they still fall short in fine-grained, font-aware text manipulation, limiting their utility in professional workflows. To address this issue, we present SkyReels-Text, a novel font-controllable framework for precise poster text editing. Our method enables simultaneous editing of multiple text regions, each rendered in distinct typographic styles, while preserving the visual appearance of non-edited regions. Notably, our model requires neither font labels nor test-time fine-tuning: users can simply provide cropped glyph patches corresponding to their desired typography - even if the font is not included in any standard library. Extensive experiments on multiple benchmarks demonstrate that SkyReels-Text achieves state-of-the-art performance in both text fidelity and visual realism, offering unprecedented control over font families and stylistic nuances. This work bridges the gap between general-purpose image editing and professional-grade typographic design. Code and models are publicly available at https://github.com/SkyworkAI/SkyReels-Text.

📄 PDF Abstract BibTeX arXiv:2511.13285

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

2025-06-01 · Zhengcong Fei, Hao Jiang, Di Qiu, Baoxuan Gu 외

The generation and editing of audio-conditioned talking portraits guided by multimodal inputs, including text, images, and videos, remains under explored. In this paper, we present SkyReels-Audio, a unified framework for…

Denoising

SkyReels-A2: Compose Anything in Video Diffusion Transformers

2025-04-03 · Zhengcong Fei, Debang Li, Di Qiu, Jiahua Wang 외

This paper presents SkyReels-A2, a controllable video generation framework capable of assembling arbitrary visual elements (e.g., characters, objects, backgrounds) into synthesized videos based on textual prompts while m…

Human-Domain Subject-to-VideoOpen-Domain Subject-to-VideoSingle-Domain Subject-to-VideoVideo Generation

FonTS: Text Rendering with Typography and Style Controls

2024-11-28 · Wenda Shi, Yiren Song, Dengming Zhang, Jiaming Liu 외

Visual text rendering are widespread in various real-world applications, requiring careful font selection and typographic choices. Recent progress in diffusion transformer (DiT)-based text-to-image (T2I) models show prom…

parameter-efficient fine-tuning

SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model

2026-02-25 · Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang 외 arxiv

SkyReels V4 is a unified multi modal video foundation model for joint video audio generation, inpainting, and editing. The model adopts a dual stream Multimodal Diffusion Transformer (MMDiT) architecture, where one branc…

Instruction FollowingAudio GenerationVideo Generation

ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations

2025-02-16 · Bowen Jiang, Yuan Yuan, Xinyi Bai, Zhuoqun Hao 외

This work demonstrates that diffusion models can achieve font-controllable multilingual text rendering using just raw images without font label annotations. Visual text rendering remains a significant challenge. While re…

Text Segmentation