paper-with-me

홈 › Papers

GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows

2026-03-12 · Zexuan Yan, Jiarui Jin, Yue Ma, Shijian Wang, Jiahui Hu, Wenxiang Jiao, Yuan Lu, Linfeng Zhang arxiv

Despite recent advances in generative models driving significant progress in text rendering, accurately generating complex text and mathematical formulas remains a formidable challenge. This difficulty primarily stems from the limited instruction-following capabilities of current models when encountering out-of-distribution prompts. To address this, we introduce GlyphBanana, alongside a corresponding benchmark specifically designed for rendering complex characters and formulas. GlyphBanana employs an agentic workflow that integrates auxiliary tools to inject glyph templates into both the latent space and attention maps, facilitating the iterative refinement of generated images. Notably, our training-free approach can be seamlessly applied to various Text-to-Image (T2I) models, achieving superior precision compared to existing baselines. Extensive experiments demonstrate the effectiveness of our proposed workflow. Associated code is publicly available at https://github.com/yuriYanZeXuan/GlyphBanana.

📄 PDF Abstract BibTeX arXiv:2603.12155

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DiffRenderGAN: Addressing Training Data Scarcity in Deep Segmentation Networks for Quantitative Nanomaterial Analysis through Differentiable Rendering and Generative Modelling

2025-02-13 · Dennis Possart, Leonid Mill, Florian Vollnhals, Tor Hildebrand 외

Nanomaterials exhibit distinctive properties governed by parameters such as size, shape, and surface characteristics, which critically influence their applications and interactions across technological, biological, and e…

Generative Adversarial Network

Advancing 3D Gaussian Splatting Editing with Complementary and Consensus Information

2025-03-14 · Xuanqi Zhang, Jieun Lee, Chris Joslin, WonSook Lee

We present a novel framework for enhancing the visual fidelity and consistency of text-guided 3D Gaussian Splatting (3DGS) editing. Existing editing approaches face two critical challenges: inconsistent geometric reconst…

3DGSDenoisingImage Manipulation

LUCAS: Layered Universal Codec Avatars

2025-02-27 · CVPR 2025 1 · Di Liu, Teng Deng, Giljoo Nam, Yu Rong 외

Photorealistic 3D head avatar reconstruction faces critical challenges in modeling dynamic face-hair interactions and achieving cross-identity generalization, particularly during expressions and head movements. We presen…

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation

2025-07-11 · Liu He, Xiao Zeng, Yizhi Song, Albert Y. C. Chen 외 arxiv

Multimodal Large Language Models (MLLMs) struggle with accurately capturing camera-object relations, especially for object orientation, camera viewpoint, and camera shots. This stems from the fact that existing MLLMs are…

Image Generation

ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation

2025-04-03 · Yuan Zhou, Shilong Jin, Litao Hua, Wanjun Lv 외

Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods leverage 3D Gaussian Splatting with score …

3D GenerationText to 3D