paper-with-me

홈 › Papers

AnyTrans: Translate AnyText in the Image with Large Scale Models

2024-06-17 · Zhipeng Qian, Pei Zhang, Baosong Yang, Kai Fan, Yiwei Ma, Derek F. Wong, Xiaoshuai Sun, Rongrong Ji

This paper introduces AnyTrans, an all-encompassing framework for the task-Translate AnyText in the Image (TATI), which includes multilingual text translation and text fusion within images. Our framework leverages the strengths of large-scale models, such as Large Language Models (LLMs) and text-guided diffusion models, to incorporate contextual cues from both textual and visual elements during translation. The few-shot learning capability of LLMs allows for the translation of fragmented texts by considering the overall context. Meanwhile, the advanced inpainting and editing abilities of diffusion models make it possible to fuse translated text seamlessly into the original image while preserving its style and realism. Additionally, our framework can be constructed entirely using open-source models and requires no training, making it highly accessible and easily expandable. To encourage advancement in the TATI task, we have meticulously compiled a test dataset called MTIT6, which consists of multilingual text image translation data from six language pairs.

📄 PDF Abstract BibTeX arXiv:2406.11432

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot LearningTranslation

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

AnyText: Multilingual Visual Text Generation And Editing

2023-11-06 · Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng 외

Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still p…

Image GenerationOptical Character Recognition (OCR)Text Generation

AnyText2: Visual Text Generation and Editing With Customizable Attributes

2024-11-22 · Yuxiang Tuo, Yifeng Geng, Liefeng Bo

As the text-to-image (T2I) domain progresses, generating text that seamlessly integrates with visual content has garnered significant attention. However, even with accurate text generation, the inability to control font …

Image GenerationText Generation

TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation

2026-08-17 · Xiaoan Liu, Lichen Ma, Zipeng Guo, Yu He 외 arxiv

Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing methods struggle to achieve accurate translat…

Reinforcement LearningImage Editing

UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

2025-07-01 · Yuanrui Wang, Cong Han, YafeiLi, Zhipeng Jin 외

Text-to-image generation has greatly advanced content creation, yet accurately rendering visual text remains a key challenge due to blurred glyphs, semantic drift, and limited style control. Existing methods often rely o…

Image GenerationText to Image GenerationText-to-Image Generation

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards

2026-05-19 · Mingxuan Cui, Jingpu Yang, Fengxian Ji, Qian Jiang 외 arxiv

Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and fine-grained glyph-level structure. Prior methods often improve this …

Instruction Following