paper-with-me

홈 › Papers

Text-to-Image Generation Grounded by Fine-Grained User Attention

2020-11-07 · Jing Yu Koh, Jason Baldridge, Honglak Lee, Yinfei Yang

Localized Narratives is a dataset with detailed natural language descriptions of images paired with mouse traces that provide a sparse, fine-grained visual grounding for phrases. We propose TReCS, a sequential model that exploits this grounding to generate images. TReCS uses descriptions to retrieve segmentation masks and predict object labels aligned with mouse traces. These alignments are used to select and position masks to generate a fully covered segmentation canvas; the final image is produced by a segmentation-to-image generator using this canvas. This multi-step, retrieval-based approach outperforms existing direct text-to-image generation models on both automatic metrics and human evaluations: overall, its generated images are more photo-realistic and better match descriptions.

📄 PDF Abstract BibTeX arXiv:2011.03775

Code (1)

google-research/trecs_image_generation

Tasks

Image GenerationPositionRetrievalSegmentationText to Image GenerationText-to-Image GenerationVisual Grounding

Similar Papers 제목 키워드 기반

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

2025-02-04 · Xueqing Deng, Qihang Yu, Ali Athar, Chenglin Yang 외

This paper introduces the COCONut-PanCap dataset, created to enhance panoptic segmentation and grounded image captioning. Building upon the COCO dataset with advanced COCONut panoptic masks, this dataset aims to overcome…

Image CaptioningPanoptic SegmentationSegmentation

TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering

2026-04-27 · Dongxing Mao, Yilin Wang, Linjie Li, Zhengyuan Yang 외 arxiv

Despite recent advances in text-to-image generation, models still struggle to accurately render prompt-specified text with correct spatial layout -- especially in multi-span, structured settings. This challenge is driven…

Text-to-Image Generation

Learning Fine-Grained Grounded Citations for Attributed Large Language Models

2024-08-08 · Lei Huang, Xiaocheng Feng, Weitao Ma, Yuxuan Gu 외

Despite the impressive performance on information-seeking tasks, large language models (LLMs) still struggle with hallucinations. Attributed LLMs, which augment generated text with in-line citations, have shown potential…

In-Context Learning

HCMA: Hierarchical Cross-model Alignment for Grounded Text-to-Image Generation

2025-05-10 · Hang Wang, Zhi-Qi Cheng, Chenhao Lin, Chao Shen 외

Text-to-image synthesis has progressed to the point where models can generate visually compelling images from natural language prompts. Yet, existing methods often fail to reconcile high-level semantic fidelity with expl…

cross-modal alignmentImage GenerationText to Image GenerationText-to-Image Generation

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding

2026-04-24 · Lihao Zheng, Zhenwei Shao, Yu Zhou, Yan Yang 외 arxiv

Although Multimodal Large Language Models (MLLMs) have advanced rapidly, they still face notable challenges in fine-grained multi-image understanding, often exhibiting spatial hallucination, attention leakage, and failur…

Image Attribution