paper-with-me

홈 › Papers

TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution

2026-01-24 · Haodong He, Xin Zhan, Yancheng Bai, Rui Lan, Lei Sun, Xiangxiang Chu arxiv

Real-world text image super-resolution aims to restore overall visual quality and text legibility in images suffering from diverse degradations and text distortions. However, the scarcity of text image data in existing datasets results in poor performance on text regions. In addition, datasets consisting of isolated text samples limit the quality of background reconstruction. To address these limitations, we construct Real-Texts, a large-scale, high-quality dataset collected from real-world images, which covers diverse scenarios and contains natural text instances in both Chinese and English. Additionally, we propose the TEXTS-Aware Diffusion Model (TEXTS-Diff) to achieve high-quality generation in both background and textual regions. This approach leverages abstract concepts to improve the understanding of textual elements within visual scenes and concrete text regions to enhance textual details. It mitigates distortions and hallucination artifacts commonly observed in text regions, while preserving high-quality visual scene fidelity. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple evaluation metrics, exhibiting superior generalization ability and text restoration accuracy in complex scenarios. All the code, model, and dataset will be released.

📄 PDF Abstract BibTeX arXiv:2601.17340

Code (0)

등록된 구현이 없습니다.

Tasks

Image Super-Resolution

Similar Papers 제목 키워드 기반

TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition

2024-12-02 · Xingsong Ye, Yongkun Du, Yunbo Tao, Zhineng Chen

Scene text recognition (STR) suffers from the challenges of either less realistic synthetic training data or the difficulty of collecting sufficient high-quality real-world data, limiting the effectiveness of trained STR…

Image GenerationOptical Character Recognition (OCR)Scene Text EditingScene Text Recognition

MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion

2023-07-03 · NeurIPS 2023 11 · Shitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang 외

This paper introduces MVDiffusion, a simple yet effective method for generating consistent multi-view images from text prompts given pixel-to-pixel correspondences (e.g., perspective crops from a panorama or multi-view i…

Image Generation

VASR: Variance-Aware Systematic Resampling for Reward-Guided Diffusion

2026-04-08 · Shivanshu Shekhar, Sagnik Mukherjee, Jia Yi Zhang, Tong Zhang arxiv

Sequential Monte Carlo (SMC) samplers for reward-guided diffusion models often suffer from rapid lineage collapse: a few high-reward particles dominate the population within a handful of resampling steps, destroying dive…

Text-to-Image Generation

Typographic Text Generation with Off-the-Shelf Diffusion Model

2024-02-22 · KhayTze Peong, Seiichi Uchida, Daichi Haraguchi

Recent diffusion-based generative models show promise in their ability to generate text images, but limitations in specifying the styles of the generated texts render them insufficient in the realm of typographic design.…

Text Generation

EmoAttack: Emotion-to-Image Diffusion Models for Emotional Backdoor Generation

2024-06-22 · Tianyu Wei, Shanmin Pang, Qi Guo, Yizhuo Ma 외

Text-to-image diffusion models can generate realistic images based on textual inputs, enabling users to convey their opinions visually through language. Meanwhile, within language, emotion plays a crucial role in express…

Backdoor AttackDiffusion PersonalizationImage Generation