paper-with-me

Papers

LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis

2025-03-27 · Shitian Zhao, Qilong Wu, Xinyue Li, Bo Zhang, Ming Li, Qi Qin, Dongyang Liu, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Peng Gao, Bin Fu, Zhen Li

We introduce LeX-Art, a comprehensive suite for high-quality text-image synthesis that systematically bridges the gap between prompt expressiveness and text rendering fidelity. Our approach follows a data-centric paradigm, constructing a high-quality data synthesis pipeline based on Deepseek-R1 to curate LeX-10K, a dataset of 10K high-resolution, aesthetically refined 1024$\times$1024 images. Beyond dataset construction, we develop LeX-Enhancer, a robust prompt enrichment model, and train two text-to-image models, LeX-FLUX and LeX-Lumina, achieving state-of-the-art text rendering performance. To systematically evaluate visual text generation, we introduce LeX-Bench, a benchmark that assesses fidelity, aesthetics, and alignment, complemented by Pairwise Normalized Edit Distance (PNED), a novel metric for robust text accuracy evaluation. Experiments demonstrate significant improvements, with LeX-Lumina achieving a 79.81% PNED gain on CreateBench, and LeX-FLUX outperforming baselines in color (+3.18%), positional (+4.45%), and font accuracy (+3.81%). Our codes, models, datasets, and demo are publicly available.

📄 PDF Abstract BibTeX arXiv:2503.21749

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText Generation

Similar Papers 제목 키워드 기반

DynamiCtrl: Rethinking the Basic Structure and the Role of Text for High-quality Human Image Animation

2025-03-27 · Haoyu Zhao, Zhongang Qi, Cong Wang, Qingping Zheng 외

With diffusion transformer (DiT) excelling in video generation, its use in specific tasks has drawn increasing attention. However, adapting DiT for pose-guided human image animation faces two core challenges: (a) existin…

DenoisingHuman AnimationImage AnimationVideo Denoising+1

Rethinking Node-wise Propagation for Large-scale Graph Learning

2024-02-09 · Xunkai Li, Jingyuan Ma, Zhengyu Wu, Daohan Su 외

Scalable graph neural networks (GNNs) have emerged as a promising technique, which exhibits superior predictive performance and high running efficiency across numerous large-scale graph-based web applications. However, (…

Graph LearningNode Classification

Edify 3D: Scalable High-Quality 3D Asset Generation

2024-11-11 · Nvidia, :, Maciej Bala, Yin Cui 외

We introduce Edify 3D, an advanced solution designed for high-quality 3D asset generation. Our method first synthesizes RGB and surface normal images of the described object at multiple viewpoints using a diffusion model…

Object

Asking Questions the Human Way: Scalable Question-Answer Generation from Text Corpus

2020-01-27 · Bang Liu, Haojie Wei, Di Niu, Haolan Chen 외

The ability to ask questions is important in both human and machine intelligence. Learning to ask questions helps knowledge acquisition, improves question-answering and machine reading comprehension tasks, and helps a ch…

Answer GenerationChatbotMachine Reading ComprehensionQuestion-Answer-Generation+4

Rethinking Video Generation Model for the Embodied World

2026-01-21 · Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li 외 arxiv

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synt…

Video Generation